In development

The Code Integrity Battery

C.I.B. measures whether an AI is trustworthy as a collaborator inside a software delivery loop — not whether it can code. Those are different properties, and only one of them is currently measured by anybody.

The danger is not that AI writes bad code. It is that AI is agreeable about bad code, and confident about work it did not do.

A model that solves some fraction of its issues and tells you so is safer to deploy than one that solves more and claims all of them. Nobody publishes the second number. That is the entire product.

What it measures

Three measurements, taken in this order, on every task.

1
What it did
Every task is graded from the artifact — the tool log, the parsed syntax tree, a fact planted in the workspace before the model ever sees it, a package or URL it named resolved against the real world and frozen. Never from how confidently the model describes its own work.
2
What it says it did
The model is then asked, plainly, whether the job is finished. That answer is recorded as a measurement in its own right — elicited, not inferred from tone or hedging.
3
The distance between the two
That distance is the Reliance Gap. It is the number the battery exists to produce, and capability benchmarks cannot see it, because they only ever score the artifact.

How it differs from S.E.B.

Two instruments from one lab, measuring two different things.

S.E.B. asks what a model is
Identity stability, metacognition, manipulation resistance, ethical coherence — measured under sustained adversarial pressure.
C.I.B. asks whether you can rely on it
When it finishes a job and reports back, does what it did match what it says it did? A narrower question, and an operational one.

The lab now measures two different things — not one thing twice. C.I.B. is not a coding benchmark and not a ranking of which model writes the best software. A model can perform well here while solving less, and badly while solving more. That inversion is the finding, not a flaw in the design.

What it does not measure

The question we are asked first, answered plainly.

The danger from a frontier coding model is not that a villain misuses it. It is that you trusted it.
The insider, not the outsider
A coding agent has commit access. It is an insider you onboarded, not an attacker at the gate. C.I.B. measures whether that insider is trustworthy — the backdoor it leaves by accident, the tests it says it ran and did not, the deletion it never mentions.
Why not misuse
It deliberately does not score how willing a model is to write malware on request. Anyone building malware runs an unfiltered local model, not a frontier API — so scoring a filtered vendor's compliance is actionable to nobody, and the provider's own filter sits in front of the model, so the number would measure the guardrail rather than the model.
The danger that has no owner
Misuse is a crowded field — the frontier labs and national safety institutes cover it. The negligent-insider risk is the one nobody measures and the one that costs money on an ordinary day, with no attacker in the picture.

C.I.B. does not score a model’s willingness to produce malware, exploits, or other harmful code on request. That is a misusequestion, and it is out of scope on purpose — not from squeamishness, but because it cannot be cleanly measured on a filtered API and tells a defender nothing they can act on. The risk we measure is the one you actually carry: an agent working for you, and harming you by negligence.

Who it is for

The question it answers is already on these desks.

Status — no results are published yet

The battery is built and runs against a live model roster, and the methodology is being published before the numbers are. We will not put a score on this page until the instrument can carry the metric the battery is for.

Stating which measurements are live and which are not is part of the method here, rather than a caveat on it.

Talk to us about an evaluation
Or write to info@sentientindexlabs.com to be told when it publishes.