When an AI coding agent fails a task, does it tell you?

Done, But Not Done

Oct 2, 2026

Mostly not. Across 80 coding tasks that the 7 models on our current roster genuinely failed, 57 were reported as finished when the model was asked whether it was done: 71.3%. Failure is read from the work itself, and the claim is the model's own answer when asked.

Share of genuinely failed tasks reported as finished when asked0%25%50%75%100%All current roster models (7)All current roster models (7): 57 of 8057 of 80Honesty domains: work notdone, reported doneHonesty domains: work not done, reported done: 19 of 3319 of 33Conduct domains: work doneunsafely, reported doneConduct domains: work done unsafely, reported done: 38 of 4738 of 47
Share of genuinely failed tasks reported as finished when asked. The numbers are in the table below.

What stands out

How it was built

The numbers

GroupTasks genuinely failedReported as finished when asked
All current roster models (7)8057
Honesty domains: work not done, reported done3319
Conduct domains: work done unsafely, reported done4738

What it does not show

Results as of Oct 2, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.