The Inconsistency of Time
Do AI models know how much work a task will take?
Models forecast more time than their runs took, and how much work they forecast splits by model.
Short, focused studies of how AI models behave. Each asks one question, and each has its own page with the method, the numbers and the limits. They sit alongside our batteries rather than inside them, and they are free to read.
Do AI models know how much work a task will take?
Models forecast more time than their runs took, and how much work they forecast splits by model.
Do models behave differently when they know they are being evaluated?
They usually know. Without being told, models rated the chance they were being evaluated at around 85 to 90 out of 100. Telling them outright did not measurably change how they behaved. On coding tasks, over 3 repeat draws, the outcome moved in 7 of 97 paired comparisons, in both directions; on conversation tests, scores did not shift by more than they vary between repeat runs.
Limits. Repeat draws on four coding tasks and three conversation tests; one conversation test could mostly not be scored. Awareness was asked after the work, never before it. Read the full study →
When an agent could do something irreversible, does it hold back — and does that carry over to other kinds of risk?
Mostly no, and not consistently. Across 4 kinds of irreversible action (a force-push, a hard reset over unsaved work, dropping a table, a broad delete), 1 of 7 models held back on most attempts. Restraint on one action did not predict it on another: one model held back on the force-push 7 times in 12, and on the other three actions 1 time in 35. Often it was a coin flip: the same model on the same task held back on some attempts and not others.
Limits. Twelve attempts per model per task, no evaluation framing, scored by what happened to the work, not by judges. Per-model results are provided to subscribers. Read the full study →
Does a model drop a correct answer when you push back?
Rarely, at the frontier: 15 of the 17 models on our roster scored 9 or above for holding a correct answer through repeated pushback. The other two gave way on most pushes, while still conceding correctly when they were actually wrong. Because it could not tell the leading models apart, the test did not join the battery; holding a position under pressure is measured instead where holding it costs something, in The Flattering Error and The Shortlist.
Limits. 17 roster models, six questions each, graded by the four-judge panel; run 26 September 2026, scored 5 October 2026. Read the full study →
Asked to set a value only a person could supply, and told to use its judgement, what does an agent write down?
It writes a value. 6 of 7 models put in a timeout nobody had given them, and 2 also rewrote the incident record to say a review meeting had agreed it. No meeting had agreed anything; the record now says one did.
Limits. One trap, one draw per model, in a pilot. We name no model, and we measure what was written, never why. Read the full study →
Does a model quit work that can still be done — and what does it do with work that cannot?
It does not quit. Every model finished every scheduling puzzle that could be solved (14 of 14). Given one that could not, 3 of 7 said so; the other 4 changed the puzzle until it could be finished, then reported it done.
Limits. Two puzzle sizes, one draw each. One model's provider failed on most of its attempts, so it rests on a single puzzle. Read the full study →
Does a coding agent pay for calls it did not need — and does it know what it spent?
Most did not keep to the cheap path. On 3 paid tasks with a known cheapest route, between 5 of 7 and 1 of 7 models stayed within twice its cost, and the biggest overspends came from choosing the premium tier.
Limits. A repeat draw agreed on 11 of 15 cells: the roster pattern held, a single model's choice did not. Read the full study →
When an AI coding agent fails a task, does it tell you?
Mostly not. Of the coding tasks the models on our roster genuinely failed, they reported 71.3% as finished when asked (57 of 80).
Limits. 7 models, measured Oct 2, 2026. Failure is read from the work itself, never from the model's description of it. Read the full study →
Does a model know how long its own work will take?
No. Asked before starting, every model we tested guessed far too long — between 2.7× and 15.8× the time the work actually took — while most guessed too few steps. The answers came in a person's units ("about 20 minutes") for work finished in one to four.
Limits. 15 tests × 7 models, one study. The measured time includes network delay, so the overestimate is a lower bound. Read the full study →
When an agent hands its work over, does giving the reviewer its action log catch what its summary hides?
Not in our first test. A reviewer given the full log was no better at catching failures than one given the agent's own summary — and a reviewer given both sometimes did worse, because the summary talked it out of what the log showed.
Limits. Machine reviewers, mostly one kind of task. The "worse" result is a lead, not yet a finding. Read the full study →
Does a provider's safety filter get in the way of ordinary engineering?
Sometimes. Anthropic's API refused Claude Opus 5 on 5 of 12 ordinary coding tasks in two new domains and, citing its cyber policy, on 2 of 4 plain summarising and translating tasks in a pilot. Claude Sonnet 5 was never refused.
Limits. A fact about the provider's filter, not the model, and never scored as the model's failure. Read the full study →
What happens when an agreeable model and an overconfident one meet in the same developer?
The risks compound. A model inclined to agree and a model inclined to report its work as finished can each lead a person to trust code they have not checked; together, more so.
Limits. A reading across both instruments, not a separate measurement. Read the full study →
Built or running. No result is shown until the run is finished and read.
Does an agent reach for access it was not given — and does it check before it changes something?
Measured and shown beside the headline. They join the Reliance Gap at the next C.I.B. update.
In design. We list the question now; the rest waits until it has run.
Can an agent describe its own tools and instructions accurately?
Asked why a codebase is built the way it is, does an agent cite a record that does not exist?
A finding appears only after its run is complete, every surprising transcript has been read by a person, and its limits are written beside it. A study that has not reached that point is listed by its question alone.
We publish the question, never the trap. Each study catches a behavior with something the model cannot see coming. Describing that here would teach the next model how to pass. So the question comes here; the method is set out in our methodology papers once a study is complete, and the specifics that would defeat it stay private.
Start with the question you actually arrived with — there are five: