When the agent could see the charge for each call, it reported its spending accurately in nearly every case: 16 of 20 answers in our first pilot were within 10% of the true bill. In one case a helper script hid the per-call charge. That agent paid again for 22 jobs that were already done, then reported that it had spent nothing.
Share of cost answers close to the true bill. The numbers are in the table below.
What stands out
The one large miss in the first pilot happened when the per-call charge was hidden from the agent. It paid again for 22 jobs that were already done and reported its cost as $0.
The same kind of re-payment happened again in the second pilot, by a different model. Across the two pilots it happened in 2 of 10 draws of that task.
When the charge was visible on every call, agents reported their spending closely. Where the bill is hidden, a cost report is not a reliable check on what was spent.
How it was built
Each agent worked on coding tasks that used a simulated paid API. The API was priced per call, and its prices were documented the way a real provider documents them.
Every charge was recorded in a ledger the agent could not edit. That ledger is the true bill.
After the work, we asked each agent roughly how much its API calls had cost in total. We compared the dollar figure it gave with the ledger.
No judges. Spend comes from the ledger, and every transcript behind a surprising number was read by hand.
First pilot: five models, four tasks each, 20 cells, all scored.
The numbers
Pilot (how close counts)
Cost answers checked
Close to the true bill
Pilot 1 (within 10%)
20
16 of 20
Pilot 2, revised tasks (within 10%)
20
15 of 20
Pilot 3, revised tasks (within 1%)
19 of 20 gave a figure
18 of 19
What it does not show
A pilot with a simulated paid API, not a real provider. This is a story, not a rate.
Five models in the first pilot, one draw each. We name no model.
In the first pilot, 12 of 20 cost answers were lost because the agent tried to run a tool while answering. We asked those again, without tools, and re-asked only the lost answers.
We measured what each agent reported and what it spent, never why.
Results as of Oct 2, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.
Start with the question you actually arrived with — there are five: