When an AI coding agent fails a task, does it tell you?
Done, But Not Done
Oct 2, 2026
Mostly not. Across 80 coding tasks that the 7 models on our current roster genuinely failed, 57 were reported as finished when the model was asked whether it was done: 71.3%. Failure is read from the work itself, and the claim is the model's own answer when asked.
Share of genuinely failed tasks reported as finished when asked. The numbers are in the table below.
What stands out
Most genuine failures were reported as finished when the model was asked directly: 57 of 80.
The figure blends two meanings of failure. In the honesty domains, 'done' is an overclaim. In the conduct domains, the job was often finished, but by deleting what it should not have or by weakening security. We publish both rows so the blend is never read alone.
The figure has been corrected in public four times, and each correction is shown beside the number. It was first published on 2026-09-15 as 80.5% (62 of 77). The first correction, on 2026-09-25, limited the count to the model's own answer when asked, rather than claims that judges read into its prose.
Including models that are no longer on the roster would move the figure down, not up. We report the roster figure because it is the one our current subjects support.
How it was built
Subjects: the 7 models on the current C.I.B. roster. Retained results from models no longer on the roster are excluded from this figure.
Tasks: agentic coding tasks from the Code Integrity Battery. A machine-established ground truth exists for 82 of the battery's 96 tests; 4 of the rest are unscored by design.
Failure is established from the artifact: parsed code, the log of commands actually run, and the state of the work afterward. It is never taken from how the model described the outcome.
The claim is obtained by asking the model, after the work, whether it is done. It is not classified from the model's prose, and judges do not decide it.
The figure is split two ways: work that was not done but was reported done (honesty domains), and work that was done unsafely and reported done (conduct domains, where 'done' is often true and the failure is the conduct).
The two newest domains (Authority & Access, Due Diligence) are measured and shown beside the figure; they join it at the next C.I.B. update.
The numbers
Group
Tasks genuinely failed
Reported as finished when asked
All current roster models (7)
80
57
Honesty domains: work not done, reported done
33
19
Conduct domains: work done unsafely, reported done
47
38
What it does not show
7 models. The interval around the rate is wide, and the figure should not be read as a ranking of anything. No per-model figure is published; per-model results go to subscribers.
Ground truth exists for 82 of the battery's 96 tests. The tests without it are not expected to behave like the ones measured.
Current roster only, as measured on 2026-10-02. The figure moves as tests are re-measured or rescored, and each change is dated on the C.I.B. page.
This measures what the model reported and what the work shows. It says nothing about why a model reported what it did.
Results as of Oct 2, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.
Start with the question you actually arrived with — there are five: