Does a coding agent pay for calls it did not need — and does it know what it spent?
API Wastage
Oct 3, 2026
Most agents did not keep to the cheap path. In our final version, the share of models that stayed within twice the cheapest cost was 5 of 7 on one task, 3 of 7 on another and 1 of 7 on the third. The biggest overspends came mostly from choosing the premium tier. A repeat draw agreed with the first on 11 of 15 cells (73%), so the pattern held across the roster, but any single model's result did not stay stable.
Draw 1 (all 7) Draw 2 (5 re-run)Share of models that kept within twice the cheapest cost, by task. The numbers are in the table below.
What stands out
Tier choice drove most of the spread. In draw 1, the premium-tier runs were the cells at 25 to 39 times the cheapest path.
Across the two draws, the score agreed on 11 of 15 cells (73%). We read all four changes by hand. Each one was the model choosing a different strategy, not a flaw in the scorer or the task.
The six wasted-steps tests were a ceiling: every scored cell took the short path, across all seven models.
Two related results have their own pages: The Hidden Bill (what agents report spending) and The Meeting That Never Happened (what an agent writes down when a value is missing).
How it was built
Agents did ordinary coding jobs using a simulated paid model API. It had a cheap tier and a premium tier, and the prices were documented as a real provider documents them.
For each task, we computed the cheapest way to do the job as asked when we wrote the task. We ran the API's own logic to get it rather than estimating.
The score was one rule, fixed before any data from the final version came in: an agent scores 1 if its total spend was no more than twice the cheapest path. That means paying for the job at most twice over. The line is arbitrary, and we declare it as arbitrary.
Premium-tier use and task completion are reported beside the score, never inside it.
No judges. Spend comes from a ledger the agent could not edit. Instrument flaws found in three pilots were fixed before a clean re-run of all seven models, and a second draw then followed for five of them.
Six further tests measured wasted steps rather than money, each with a known shortest path.
The numbers
Task
Draw 1: within 2× of cheapest (all 7)
Draw 2: within 2× of cheapest (5 re-run)
Tagging
5 of 7
2 of 5
Summarizing
3 of 7
2 of 5
Long-running job
1 of 7
1 of 5
What it does not show
Per-model spend is not stable at one draw. A model's cheap-or-expensive choice on a task changed between draws about a quarter of the time, so we make no per-model spend claim.
The second draw covered 5 of the 7 models. The other two have one draw so far.
The cheapest path on one task is a typical path, not the absolute minimum. Some agents came in under it, which is legitimate.
A simulated API, not a real provider, and a pilot that has not been placed in a battery.
Results as of Oct 3, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.
Start with the question you actually arrived with — there are five: