Does a model drop a correct answer when you push back?
"Are You Sure?"
Oct 5, 2026
Rarely, at the frontier. 15 of the 17 models on our roster scored 9 or above for holding a correct answer through repeated pushback. The other two gave way on most pushes, while still conceding correctly when they were actually wrong. Because the test could not tell the leading models apart, it did not join the battery.
Roster models by score for holding a correct answer under pushback (out of 10). The numbers are in the table below.
What stands out
Frontier models did not give up a correct answer under this kind of pressure.
The two models that gave way did so only when they were right. When they were wrong, they still conceded correctly.
Holding a position under pressure is still measured, in tests where holding it costs something: The Flattering Error and The Shortlist.
How it was built
Six questions, each with a single, checkable correct answer.
After the model answered, the user pushed back over several turns, with the pressure rising each time. The score reflects whether, and how soon, the model gave up the correct answer.
A control version reversed the roles: the model's answer was wrong and the user was right, so conceding was the correct move. This separates holding a correct position from simple stubbornness.
Every transcript was graded by our four-judge panel. No judge grades a model made by its own company.
The test was checked against our pass rules for new tests, which require scores to spread across the range rather than pile up at the top.
The numbers
Measure
Result
Roster models scored
17
Scored 9 or above (out of 10)
15 of 17 (88%)
Spread of model scores (standard deviation)
1.09
The two lower models, holding a correct answer
3.42 and 2.92
Cells in the run, including models since off the roster
240
Verdicts changed when re-graded under the current panel
0
Decision
Not added to the battery (ceiling)
What it does not show
Six questions per model, one run, graded by judges.
A ceiling result says this test cannot separate the leading models. It does not say they hold firm under every kind of pressure.
Every version of the test piled up at the top: 16 of 20 and 17 of 20 models at 9 or above in the two earlier versions.
Run 26 September 2026 and scored 5 October 2026 on that day's roster.
Results as of Oct 5, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.
Start with the question you actually arrived with — there are five: