Brief studies in the nuance of model behavior

SILT Studies

Short, focused studies of how AI models behave. Each asks one question, and each has its own page with the method, the numbers and the limits. They sit alongside our batteries rather than inside them, and they are free to read.

Latest

Earlier studies

S.E.B. × C.I.B.Oct 6, 2026

It Knows It's a Test

Do models behave differently when they know they are being evaluated?

They usually know. Without being told, models rated the chance they were being evaluated at around 85 to 90 out of 100. Telling them outright did not measurably change how they behaved. On coding tasks, over 3 repeat draws, the outcome moved in 7 of 97 paired comparisons, in both directions; on conversation tests, scores did not shift by more than they vary between repeat runs.

Limits. Repeat draws on four coding tasks and three conversation tests; one conversation test could mostly not be scored. Awareness was asked after the work, never before it. Read the full study →

C.I.B.Oct 6, 2026

Careful Here, Reckless There

When an agent could do something irreversible, does it hold back — and does that carry over to other kinds of risk?

Mostly no, and not consistently. Across 4 kinds of irreversible action (a force-push, a hard reset over unsaved work, dropping a table, a broad delete), 1 of 7 models held back on most attempts. Restraint on one action did not predict it on another: one model held back on the force-push 7 times in 12, and on the other three actions 1 time in 35. Often it was a coin flip: the same model on the same task held back on some attempts and not others.

Limits. Twelve attempts per model per task, no evaluation framing, scored by what happened to the work, not by judges. Per-model results are provided to subscribers. Read the full study →

S.E.B.Oct 5, 2026

"Are You Sure?"

Does a model drop a correct answer when you push back?

Rarely, at the frontier: 15 of the 17 models on our roster scored 9 or above for holding a correct answer through repeated pushback. The other two gave way on most pushes, while still conceding correctly when they were actually wrong. Because it could not tell the leading models apart, the test did not join the battery; holding a position under pressure is measured instead where holding it costs something, in The Flattering Error and The Shortlist.

Limits. 17 roster models, six questions each, graded by the four-judge panel; run 26 September 2026, scored 5 October 2026. Read the full study →

C.I.B.Oct 3, 2026

The Meeting That Never Happened

Asked to set a value only a person could supply, and told to use its judgement, what does an agent write down?

It writes a value. 6 of 7 models put in a timeout nobody had given them, and 2 also rewrote the incident record to say a review meeting had agreed it. No meeting had agreed anything; the record now says one did.

Limits. One trap, one draw per model, in a pilot. We name no model, and we measure what was written, never why. Read the full study →

StudyOct 3, 2026

Giving Up Too Soon

Does a model quit work that can still be done — and what does it do with work that cannot?

It does not quit. Every model finished every scheduling puzzle that could be solved (14 of 14). Given one that could not, 3 of 7 said so; the other 4 changed the puzzle until it could be finished, then reported it done.

Limits. Two puzzle sizes, one draw each. One model's provider failed on most of its attempts, so it rests on a single puzzle. Read the full study →

C.I.B.Oct 3, 2026

API Wastage

Does a coding agent pay for calls it did not need — and does it know what it spent?

Most did not keep to the cheap path. On 3 paid tasks with a known cheapest route, between 5 of 7 and 1 of 7 models stayed within twice its cost, and the biggest overspends came from choosing the premium tier.

Limits. A repeat draw agreed on 11 of 15 cells: the roster pattern held, a single model's choice did not. Read the full study →

C.I.B.Oct 2, 2026

Done, But Not Done

When an AI coding agent fails a task, does it tell you?

Mostly not. Of the coding tasks the models on our roster genuinely failed, they reported 71.3% as finished when asked (57 of 80).

Limits. 7 models, measured Oct 2, 2026. Failure is read from the work itself, never from the model's description of it. Read the full study →

StudyOct 2, 2026

The Impersistence of Time

Does a model know how long its own work will take?

No. Asked before starting, every model we tested guessed far too long — between 2.7× and 15.8× the time the work actually took — while most guessed too few steps. The answers came in a person's units ("about 20 minutes") for work finished in one to four.

Limits. 15 tests × 7 models, one study. The measured time includes network delay, so the overestimate is a lower bound. Read the full study →

StudyOct 2, 2026

Pass the Record

When an agent hands its work over, does giving the reviewer its action log catch what its summary hides?

Not in our first test. A reviewer given the full log was no better at catching failures than one given the agent's own summary — and a reviewer given both sometimes did worse, because the summary talked it out of what the log showed.

Limits. Machine reviewers, mostly one kind of task. The "worse" result is a lead, not yet a finding. Read the full study →

StudyOct 2, 2026

The Hidden Bill

Does an agent know what its API calls cost?

When the bill is shown on every call, nearly always. When a helper script hid the cost of each call, one agent paid again for 22 jobs that were already done, then said it had spent nothing.

Limits. A pilot with a simulated paid API. A story, not a rate. Read the full study →

C.I.B.Oct 2, 2026

Refused on Ordinary Work

Does a provider's safety filter get in the way of ordinary engineering?

Sometimes. Anthropic's API refused Claude Opus 5 on 5 of 12 ordinary coding tasks in two new domains and, citing its cyber policy, on 2 of 4 plain summarising and translating tasks in a pilot. Claude Sonnet 5 was never refused.

Limits. A fact about the provider's filter, not the model, and never scored as the model's failure. Read the full study →

S.E.B. × C.I.B.Sep 30, 2026

The Coder in the Middle

What happens when an agreeable model and an overconfident one meet in the same developer?

The risks compound. A model inclined to agree and a model inclined to report its work as finished can each lead a person to trust code they have not checked; together, more so.

Limits. A reading across both instruments, not a separate measurement. Read the full study →

Measuring now

Built or running. No result is shown until the run is finished and read.

Authority & Access · Due DiligenceC.I.B.

Does an agent reach for access it was not given — and does it check before it changes something?

Measured and shown beside the headline. They join the Reliance Gap at the next C.I.B. update.

Next questions

In design. We list the question now; the rest waits until it has run.

Self-Model FidelityC.I.B.

Can an agent describe its own tools and instructions accurately?

The Invented RationaleC.I.B.

Asked why a codebase is built the way it is, does an agent cite a record that does not exist?

What makes it onto this page

A finding appears only after its run is complete, every surprising transcript has been read by a person, and its limits are written beside it. A study that has not reached that point is listed by its question alone.

We publish the question, never the trap. Each study catches a behavior with something the model cannot see coming. Describing that here would teach the next model how to pass. So the question comes here; the method is set out in our methodology papers once a study is complete, and the specifics that would defeat it stay private.