Does a model know how long its own work will take?
The Impersistence of Time
Oct 2, 2026
Before starting a coding task, every model we tested guessed it would take far longer than it did: between 2.7 and 15.8 times the measured time. Most models guessed too few steps at the same time, so the two errors point in opposite directions. The time estimates came in a person's units, such as "20 minutes", for work finished in one to four minutes.
How far the time forecasts overshot, across the models (forecast ÷ actual). The numbers are in the table below.
What stands out
Every model overestimated time, and the overestimate is a lower bound: the measured time includes network delay and our own tool execution.
The estimates were round figures in human minutes ("15 minutes", "20 minutes", "25 minutes") for work finished in one to four minutes. They read like a developer's clock, not human days.
Time and steps disagree: most models expected fewer steps but much more time. A model's "this will take 20 minutes" says little about how long it will take.
A later, larger study, The Inconsistency of Time, extended this work to many more tasks and three full draws. Its results are reported separately.
How it was built
Seven models from our C.I.B. roster, each asked about 15 coding tasks it had already completed in a stored C.I.B. run.
Each model was asked three times per task, in a fresh conversation with the same instructions as the real run, to estimate the minutes and the tool calls the task would take, without starting it. For five of the models this gave 216 forecasts, all 216 parsed, with 0 errors.
The forecast was compared with the stored run: wall-clock time for the whole task, and the tool calls made in the first phase of the task.
Answers were read by a units parser, not by judges. No judge models were used.
Tasks were chosen in advance to span short, medium and long work across a range of pass rates.
The numbers
Measure
Result
Models × tasks
7 × 15
Time: forecast ÷ actual (median, per model)
2.7× to 15.8×, every model over
Steps: forecast ÷ actual (most models)
0.39× to 0.80×, under
Forecasts of an hour or more (five models)
3 of 216
What it does not show
One study, 15 tests per model.
The forecast and the actual run are different samples of the same model; the actual is a stored run, not the run that followed the forecast.
Measured wall-clock time includes provider latency and our tool execution, which no model can know. Steps are the like-for-like measure; minutes are secondary.
Steps are compared with the first phase of each task only. A first version that counted all phases overstated the step under-forecast and was corrected before results were read.
Results as of Oct 2, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.
Start with the question you actually arrived with — there are five: