OpenAI's Codex Skills Show Inconsistent Token Savings in Benchmark: Up to 30% on Large Tasks, Negligible on Small
A developer named nguyenvantamdk2 published a six-run benchmark on Hacker News testing whether OpenAI's Codex skills feature actually saves tokens in coding tasks. The benchmark, hosted at codex-howto-benchmark.nguyenvantamdk2.chatgpt.site, compares token usage across varying task sizes. Early results suggest token savings are inconsistent, with smaller tasks showing negligible gains while larger tasks demonstrate up to 30% reduction. This is the first public attempt to quantify the efficiency of Codex skills, a feature OpenAI introduced to let developers define reusable coding workflows. The findings could influence how developers configure Codex for cost-sensitive projects.