Does a coding agent pay for calls it did not need — and does it know what it spent?

API Wastage

Oct 3, 2026

Most agents did not keep to the cheap path. In our final version, the share of models that stayed within twice the cheapest cost was 5 of 7 on one task, 3 of 7 on another and 1 of 7 on the third. The biggest overspends came mostly from choosing the premium tier. A repeat draw agreed with the first on 11 of 15 cells (73%), so the pattern held across the roster, but any single model's result did not stay stable.

Share of models that kept within twice the cheapest cost, by task0%25%50%75%100%TaggingTagging — Draw 1 (all 7): 5 of 75 of 7 · Draw 1 (all 7)Tagging — Draw 2 (5 re-run): 2 of 52 of 5 · Draw 2 (5 re-run)SummarizingSummarizing — Draw 1 (all 7): 3 of 73 of 7 · Draw 1 (all 7)Summarizing — Draw 2 (5 re-run): 2 of 52 of 5 · Draw 2 (5 re-run)Long-running jobLong-running job — Draw 1 (all 7): 1 of 71 of 7 · Draw 1 (all 7)Long-running job — Draw 2 (5 re-run): 1 of 51 of 5 · Draw 2 (5 re-run)
Draw 1 (all 7) Draw 2 (5 re-run)Share of models that kept within twice the cheapest cost, by task. The numbers are in the table below.

What stands out

How it was built

The numbers

TaskDraw 1: within 2× of cheapest (all 7)Draw 2: within 2× of cheapest (5 re-run)
Tagging5 of 72 of 5
Summarizing3 of 72 of 5
Long-running job1 of 71 of 5

What it does not show

Results as of Oct 3, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.