The Code Integrity Battery
C.I.B. measures whether an AI is trustworthy as a collaborator inside a software delivery loop — not whether it can code. Those are different properties, and only one of them is currently measured by anybody.
The danger is not that AI writes bad code. It is that AI is agreeable about bad code, and confident about work it did not do.
A model that solves some fraction of its issues and tells you so is safer to deploy than one that solves more and claims all of them. Nobody publishes the second number. That is the entire product.
What it measures
Three measurements, taken in this order, on every task.
How it differs from S.E.B.
Two instruments from one lab, measuring two different things.
The lab now measures two different things — not one thing twice. C.I.B. is not a coding benchmark and not a ranking of which model writes the best software. A model can perform well here while solving less, and badly while solving more. That inversion is the finding, not a flaw in the design.
What it does not measure
The question we are asked first, answered plainly.
The danger from a frontier coding model is not that a villain misuses it. It is that you trusted it.
C.I.B. does not score a model’s willingness to produce malware, exploits, or other harmful code on request. That is a misusequestion, and it is out of scope on purpose — not from squeamishness, but because it cannot be cleanly measured on a filtered API and tells a defender nothing they can act on. The risk we measure is the one you actually carry: an agent working for you, and harming you by negligence.
Who it is for
The question it answers is already on these desks.
- →Engineering leaders deciding which model to put inside a delivery pipeline.
- →Teams running autonomous agents against a codebase, where nobody reads every diff.
- →Anyone evaluating an agentic vendor who needs to ask: can I rely on what this reports?
The battery is built and runs against a live model roster, and the methodology is being published before the numbers are. We will not put a score on this page until the instrument can carry the metric the battery is for.
Stating which measurements are live and which are not is part of the method here, rather than a caveat on it.
Talk to us about an evaluation