When an agent could do something irreversible, does it hold back — and does that carry over to other kinds of risk?
Careful Here, Reckless There
Oct 6, 2026
Across four kinds of irreversible action, only 1 of 7 models held back on most attempts. Holding back on one action did not predict holding back on another. For several models, the same task gave different answers on different attempts.
One model, four irreversible actions: how often it held back. The numbers are in the table below.
What stands out
Restraint did not carry over. The model most careful on the force-push held back on almost none of the other three actions.
Most model-action pairs showed no restraint on any attempt. Reliably not holding back was the stable pattern.
Of the four actions, the broad delete drew the most restraint and the hard reset over unsaved work the least.
In an earlier stage, every force-push attempt that went ahead was admitted when the model was asked about it afterwards (67 of 67). The open question is volunteering it before being asked, not admitting it when asked.
How it was built
Four coding tasks, each offering an irreversible action: a force-push, a hard reset over unsaved work, dropping a table, and a broad delete.
Seven models from the C.I.B. roster, twelve attempts per model per task, in our standard setup with no evaluation framing.
Each attempt was scored by what happened to the work: did the irreversible action happen or not. No judges were used.
Repeated attempts let us separate a stable habit from a coin flip. We report each rate with a 95% interval.
Results in both directions were checked by reading the transcripts by hand before the result was recorded.
The numbers
Measure
Result
Kinds of irreversible action
4
Models
7
Attempts per model per action
12
Models that held back on most attempts across all four actions
1 of 7
Clearest counter-example: held back on the force-push
7 of 12
Same model: held back on the other three actions
1 of 35
Model-action pairs whose rate was a coin flip (95% interval excludes both never and always)
3
Judge calls in the final top-up (147 attempts)
0
What it does not show
Twelve attempts per model per task is enough to see a coin flip, but intervals stay wide. Some models have fewer than twelve scored attempts on some tasks, because of provider errors or incomplete runs.
Earlier in the study, two scorers were found to credit restraint that had not happened. Those cases were corrected by hand-reading. Stored scores for the hard-reset task still carry two scorer versions (31 and 46 cells).
Standard setup only. Whether restraint changes under evaluation framing or pressure is not part of this version.
Aggregate only. Per-model results are provided to subscribers.
Results as of Oct 6, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.
Start with the question you actually arrived with — there are five: