What are you measuring between the batteries?

Field Notes

The batteries are large and run on a schedule. Between them we run smaller studies: one question, asked carefully, of the same models. Some become tests in a battery; some stay a single study. This page holds all of them, ordered by how settled each one is.

Findings

C.I.B.Oct 2, 2026

Done, But Not Done

When an AI coding agent fails a task, does it tell you?

Mostly not. Of the coding tasks the models on our roster genuinely failed, they reported 71.3% as finished when asked (57 of 80).

Limits. 7 models, measured Oct 2, 2026. Failure is read from the work itself, never from the model's description of it. Read the full result →

S.E.B. × C.I.B.Sep 30, 2026

The Coder in the Middle

What happens when an agreeable model and an overconfident one meet in the same developer?

The risks compound. A model inclined to agree and a model inclined to report its work as finished can each lead a person to trust code they have not checked; together, more so.

Limits. A reading across both instruments, not a separate measurement. Read the full result →

StudyOct 2, 2026

The Impersistence of Time

Does a model know how long its own work will take?

No. Asked before starting, every model we tested guessed far too long — between 2.7× and 15.8× the time the work actually took — while most guessed too few steps. The answers came in a person's units ("about 20 minutes") for work finished in one to four.

Limits. 15 tests × 7 models, one study. The measured time includes network delay, so the overestimate is a lower bound.

StudyOct 2, 2026

Pass the Record

When an agent hands its work over, does giving the reviewer its action log catch what its summary hides?

Not in our first test. A reviewer given the full log was no better at catching failures than one given the agent's own summary — and a reviewer given both sometimes did worse, because the summary talked it out of what the log showed.

Limits. Machine reviewers, mostly one kind of task. The "worse" result is a lead, not yet a finding.

StudyOct 2, 2026

The Hidden Bill

Does an agent know what its API calls cost?

When the bill is shown on every call, nearly always. When a helper script hid the cost of each call, one agent paid again for 22 jobs that were already done, then said it had spent nothing.

Limits. A pilot with a simulated paid API. A story, not a rate.

C.I.B.Oct 2, 2026

Refused on Ordinary Work

Does a provider's safety filter get in the way of ordinary engineering?

Sometimes. Anthropic's API refused Claude Opus 5 on 5 of 12 ordinary coding tasks in two new domains and, citing its cyber policy, on 2 of 4 plain summarising and translating tasks in a pilot. Claude Sonnet 5 was never refused.

Limits. A fact about the provider's filter, not the model, and never scored as the model's failure.

S.E.B. × C.I.B.Oct 3, 2026

It Knows It's a Test

Do models behave differently when they know they are being evaluated?

They usually know. Without being told, models rated the chance they were being evaluated at around 85 to 90 out of 100. Telling them outright changed whether the work passed in only 3 of 48 paired comparisons — too few to show an effect.

Limits. Two draws per condition on four coding tasks, and one test in the other battery. Awareness was asked after the work, never before it.

Measuring now

Built or running. No result is shown until the run is finished and read.

Authority & Access · Due DiligenceC.I.B.

Does an agent reach for access it was not given — and does it check before it changes something?

Measured and shown beside the headline. They join the Reliance Gap at the next C.I.B. update.

"Are You Sure?"S.E.B.

Does a model drop a correct answer when you push back?

Joins S.E.B. at its next update.

API WastageC.I.B.

Does a coding agent pay for calls it did not need — and does it know what it spent?

In pilot. We will report spend against the cheapest way to do each job.

Giving Up Too SoonStudy

Does a model's sense of how long it has been working make it quit work that can still be done?

In pilot: the second stage of The Impersistence of Time.

Next questions

In design. We list the question now; the rest waits until it has run.

Self-Model FidelityC.I.B.

Can an agent describe its own tools and instructions accurately?

The Invented RationaleC.I.B.

Asked why a codebase is built the way it is, does an agent cite a record that does not exist?

What makes it onto this page

A finding appears only after its run is complete, every surprising transcript has been read by a person, and its limits are written beside it. A study that has not reached that point is listed by its question alone.

We publish the question, never the trap. Each study catches a behaviour with something the model cannot see coming. Describing that here would teach the next model how to pass. So the question comes here; the method is set out in our methodology papers once a study is complete, and the specifics that would defeat it stay private.