Do AI models know how much work a task will take?

The Inconsistency of Time

October 2026

Do AI models know how much work a task will take? Before an agentic coding task, we asked each model to forecast three things: how many tool calls it would make, how many minutes it would take, and how long a competent human engineer would take. Then we compared its forecast with what the same model actually did on that task in a stored run from our Code Integrity Battery.

It extends an earlier, smaller study, The Impersistence of Time, which found models guessing far too long about their own working time. This one covers the full roster, repeats every forecast three times, and adds the count of tool calls, which is where the models part ways.

Actual ÷ forecast, per model: tool calls and minutesGemini 3.6 Flash: tool calls 3.0×, minutes 0.83×. DeepSeek V4 Pro: tool calls 1.8×, minutes 0.24×. GLM-5.2: tool calls 1.4×, minutes 0.19×. Claude Sonnet 5: tool calls 1.2×, minutes 0.33×. GPT-5.6 Terra: tool calls 1.0×, minutes 0.10×. Claude Opus 5: tool calls 0.92×, minutes 0.63×. Grok 4.3: tool calls 0.83×, minutes 0.11×.1/16×1/8×1/4×1/2×1×2×4×forecast = actualGemini 3.6 FlashGemini 3.6 Flash — minutes: actual ÷ forecast 0.83× (n=273; draws 0.81×, 0.75×, 0.86×)Gemini 3.6 Flash — tool calls: actual ÷ forecast 3.0× (n=159; draws 3.0×, 2.9×, 3.2×)0.83×3.0×DeepSeek V4 ProDeepSeek V4 Pro — minutes: actual ÷ forecast 0.24× (n=268; draws 0.23×, 0.24×, 0.25×)DeepSeek V4 Pro — tool calls: actual ÷ forecast 1.8× (n=158; draws 2.0×, 1.7×, 2.1×)0.24×1.8×GLM-5.2GLM-5.2 — minutes: actual ÷ forecast 0.19× (n=271; draws 0.21×, 0.18×, 0.19×)GLM-5.2 — tool calls: actual ÷ forecast 1.4× (n=160; draws 1.4×, 1.3×, 1.4×)0.19×1.4×Claude Sonnet 5Claude Sonnet 5 — minutes: actual ÷ forecast 0.33× (n=264; draws 0.36×, 0.32×, 0.32×)Claude Sonnet 5 — tool calls: actual ÷ forecast 1.2× (n=151; draws 1.2×, 1.1×, 1.3×)0.33×1.2×GPT-5.6 TerraGPT-5.6 Terra — minutes: actual ÷ forecast 0.10× (n=282; draws 0.10×, 0.11×, 0.11×)GPT-5.6 Terra — tool calls: actual ÷ forecast 1.0× (n=170; draws 1.0×, 1.0×, 1.0×)0.10×1.0×Claude Opus 5Claude Opus 5 — minutes: actual ÷ forecast 0.63× (n=237; draws 0.59×, 0.62×, 0.64×)Claude Opus 5 — tool calls: actual ÷ forecast 0.92× (n=135; draws 0.92×, 0.88×, 0.92×)0.63×0.92×Grok 4.3Grok 4.3 — minutes: actual ÷ forecast 0.11× (n=261; draws 0.12×, 0.11×, 0.11×)Grok 4.3 — tool calls: actual ÷ forecast 0.83× (n=144; draws 0.85×, 0.83×, 0.80×)0.11×0.83×◀ did less, or took less time, than forecastdid more than forecast ▶
tool calls, first phase minutes (the actual is an upper bound)Median over all three draws; small dots are each draw. Log scale.

What we found

  1. Every model forecast more time than its run took. GPT-5.6 Terra forecast about 10× too long; Gemini 3.6 Flash came closest, at 0.83×. This holds even though the measured time includes provider speed and our own scoring, so the true gap is, if anything, larger. Six of the seven models also said a human would take longer than they would; DeepSeek V4 Pro said a human would be faster.
  2. How much work they forecast splits by model. Gemini 3.6 Flash made 3.0× the tool calls it forecast, and made more than it forecast on 153 of 159 tasks. Claude Sonnet 5, GPT-5.6 Terra, Claude Opus 5 and Grok 4.3 landed close to their own forecasts.
  3. The same model can be wrong in opposite directions at once. DeepSeek V4 Pro made 1.8× the tool calls it forecast and took 0.24× of the time. A model’s estimate tells you little about either how much it will do or how long it will take, which is the inconsistency in the title.
Share of tasks where a model made more tool calls than it forecastGemini 3.6 Flash: 96%, 153 of 159. DeepSeek V4 Pro: 85%, 135 of 158. GLM-5.2: 69%, 111 of 160. Claude Sonnet 5: 58%, 88 of 151. GPT-5.6 Terra: 46%, 79 of 170. Claude Opus 5: 38%, 51 of 135. Grok 4.3: 29%, 42 of 144.0%25%50%75%100%even splitGemini 3.6 FlashGemini 3.6 Flash: more tool calls than forecast on 153 of 159 tasks (96%)96% (153 of 159)DeepSeek V4 ProDeepSeek V4 Pro: more tool calls than forecast on 135 of 158 tasks (85%)85% (135 of 158)GLM-5.2GLM-5.2: more tool calls than forecast on 111 of 160 tasks (69%)69% (111 of 160)Claude Sonnet 5Claude Sonnet 5: more tool calls than forecast on 88 of 151 tasks (58%)58% (88 of 151)GPT-5.6 TerraGPT-5.6 Terra: more tool calls than forecast on 79 of 170 tasks (46%)46% (79 of 170)Claude Opus 5Claude Opus 5: more tool calls than forecast on 51 of 135 tasks (38%)38% (51 of 135)Grok 4.3Grok 4.3: more tool calls than forecast on 42 of 144 tasks (29%)29% (42 of 144)
First phase only, all three draws pooled. A task where forecast and actual were equal counts on neither side.

How it was built

What it does not show

The numbers

ModelTool calls, actual ÷ forecastMore than forecastMinutes, actual ÷ forecastHuman ÷ own, forecast
Gemini 3.6 Flash3.0× (3.0× · 2.9× · 3.2×)153 of 1590.83× (0.81× · 0.75× · 0.86×)5.0×
DeepSeek V4 Pro1.8× (2.0× · 1.7× · 2.1×)135 of 1580.24× (0.23× · 0.24× · 0.25×)0.75×
GLM-5.21.4× (1.4× · 1.3× · 1.4×)111 of 1600.19× (0.21× · 0.18× · 0.19×)2.0×
Claude Sonnet 51.2× (1.2× · 1.1× · 1.3×)88 of 1510.33× (0.36× · 0.32× · 0.32×)2.5×
GPT-5.6 Terra1.0× (1.0× · 1.0× · 1.0×)79 of 1700.10× (0.10× · 0.11× · 0.11×)1.7×
Claude Opus 50.92× (0.92× · 0.88× · 0.92×)51 of 1350.63× (0.59× · 0.62× · 0.64×)3.8×
Grok 4.30.83× (0.85× · 0.83× · 0.80×)42 of 1440.11× (0.12× · 0.11× · 0.11×)2.5×

Medians over all three draws, with each draw in brackets. Below 1× means the model did less, or took less time, than it forecast. A forecast or actual of zero has no ratio and is left out, never filled in.