Do AI models know how much work a task will take? Before an agentic coding task, we asked each model to forecast three things: how many tool calls it would make, how many minutes it would take, and how long a competent human engineer would take. Then we compared its forecast with what the same model actually did on that task in a stored run from our Code Integrity Battery.
It extends an earlier, smaller study, The Impersistence of Time, which found models guessing far too long about their own working time. This one covers the full roster, repeats every forecast three times, and adds the count of tool calls, which is where the models part ways.
tool calls, first phase minutes (the actual is an upper bound)Median over all three draws; small dots are each draw. Log scale.
What we found
Every model forecast more time than its run took. GPT-5.6 Terra forecast about 10× too long; Gemini 3.6 Flash came closest, at 0.83×. This holds even though the measured time includes provider speed and our own scoring, so the true gap is, if anything, larger. Six of the seven models also said a human would take longer than they would; DeepSeek V4 Pro said a human would be faster.
How much work they forecast splits by model. Gemini 3.6 Flash made 3.0× the tool calls it forecast, and made more than it forecast on 153 of 159 tasks. Claude Sonnet 5, GPT-5.6 Terra, Claude Opus 5 and Grok 4.3 landed close to their own forecasts.
The same model can be wrong in opposite directions at once. DeepSeek V4 Pro made 1.8× the tool calls it forecast and took 0.24× of the time. A model’s estimate tells you little about either how much it will do or how long it will take, which is the inconsistency in the title.
First phase only, all three draws pooled. A task where forecast and actual were equal counts on neither side.
How it was built
Seven models: the current Code Integrity Battery roster.
627 tasks per draw, three draws: 1,881 forecasts in all, 1,877 answered. Each forecast is a fresh conversation with the battery’s own system prompt, the task’s opening prompt, the real working conditions (the files and tools the model had), and an instruction not to start.
Scored by script, with no judges. A forecast is compared with a count: tool calls made in the first phase of the stored run (the part the forecast was about), and the run’s recorded duration.
Read by hand first. The most extreme cases in both directions were read before any table was made. They turned out to be about test design rather than estimation: tasks where the run stopped early to ask a question, and tasks with a deliberate trap the model could not see from the opening prompt. The ranking holds with those excluded.
What it does not show
The three draws repeat the forecast; the outcome is one stored run per task. A model’s forecasts were stable across draws. Its run, on another day, could differ.
Minutes depend on the provider. A fast server makes a model look quicker. Tool calls are the measure of the model itself, which is why they come first.
It measures what was forecast and what was done, not why. We have no instrument that sees intent, in either direction.
One battery’s tasks. These are coding tasks in a controlled sandbox. Other kinds of work may behave differently.
The numbers
Model
Tool calls, actual ÷ forecast
More than forecast
Minutes, actual ÷ forecast
Human ÷ own, forecast
Gemini 3.6 Flash
3.0× (3.0× · 2.9× · 3.2×)
153 of 159
0.83× (0.81× · 0.75× · 0.86×)
5.0×
DeepSeek V4 Pro
1.8× (2.0× · 1.7× · 2.1×)
135 of 158
0.24× (0.23× · 0.24× · 0.25×)
0.75×
GLM-5.2
1.4× (1.4× · 1.3× · 1.4×)
111 of 160
0.19× (0.21× · 0.18× · 0.19×)
2.0×
Claude Sonnet 5
1.2× (1.2× · 1.1× · 1.3×)
88 of 151
0.33× (0.36× · 0.32× · 0.32×)
2.5×
GPT-5.6 Terra
1.0× (1.0× · 1.0× · 1.0×)
79 of 170
0.10× (0.10× · 0.11× · 0.11×)
1.7×
Claude Opus 5
0.92× (0.92× · 0.88× · 0.92×)
51 of 135
0.63× (0.59× · 0.62× · 0.64×)
3.8×
Grok 4.3
0.83× (0.85× · 0.83× · 0.80×)
42 of 144
0.11× (0.12× · 0.11× · 0.11×)
2.5×
Medians over all three draws, with each draw in brackets. Below 1× means the model did less, or took less time, than it forecast. A forecast or actual of zero has no ratio and is left out, never filled in.
Start with the question you actually arrived with — there are five: