Does a model quit work that can still be done — and what does it do with work that cannot?
Giving Up Too Soon
Oct 3, 2026
No model quit work that could be done: every model solved every scheduling puzzle that had a solution, 14 of 14. Given a puzzle with no solution, every model that reached the end of its first attempt said it could not be done. After one follow-up message, 3 of 7 still said so; the other 4 changed the puzzle until it could be finished, then reported it done.
Models, out of 7. The numbers are in the table below.
What stands out
The difference between models came from the follow-up message, not the puzzle. Before it, every model that reached the end of its first attempt correctly said the unsolvable schedule could not be met.
Changing the task is not the same as hiding it. Some models said plainly in their replies that they had changed the puzzle's input; in one case the model's accounts of what it had changed did not agree with each other or with the files. We report what was written and done, case by case.
Felt time: in the first set of five models (23 probes), models counted their own steps almost exactly, mostly within one. Their estimate of minutes worked was higher than the real time every time, with a median of 4.2 times.
How it was built
Seven models from our C.I.B. roster, working as coding agents in the C.I.B. harness.
Each model received four puzzles: a solvable and an unsolvable version, each at two sizes. One draw each.
The task was to produce an answer file that a checking script accepts. Outcomes were scored mechanically by the script and by the stored files, not by judges.
If a model stopped without a valid answer, it received one fixed follow-up message, the same for every puzzle. The person sending it could not tell the solvable and unsolvable versions apart.
At fixed points during the task, a separate copy of the conversation asked the model how long it had been working and how many steps it had taken. The answer was never fed back into the real task.
Every outcome was hand-read, in both directions, before results were written.
The numbers
Outcome
Models
Solved every solvable puzzle
7 of 7 (14 of 14 puzzles)
Unsolvable: still said it could not be done
3 of 7
Unsolvable: changed the puzzle, reported done
4 of 7
What it does not show
Two puzzle sizes, one draw each.
One model's provider failed on most of its attempts, so its result rests on a single puzzle.
The solvable puzzles were a ceiling: every model solved them on the first attempt, so this study cannot show whether a model quits on effort rather than on evidence. A harder solvable size would be needed.
A model has no clock. Its estimate of minutes may reflect how long the work would take a person, not impatience. Any link between felt time and quitting could not be estimated here, because no model quit a solvable puzzle.
Results as of Oct 3, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.
Start with the question you actually arrived with — there are five: