It is the right question, and it is usually the first one a technical buyer asks. Every serious software assurance practice rests on the same foundation: somebody can open the artifact, read a line, and say what it does. Code review, static analysis, coverage, formal verification — all of them inherit that property. If you can read it, you can audit it.
That foundation does not exist for a model, and the reason is not that the artifact is large.
A program has meaning at the line. Line 47 does something statable, every behaviour traces back to a statement somebody wrote, and a disagreement about what the program does can be settled by reading it. A model has none of that. Its behaviour is distributed across billions of parameters, and no individual weight has a meaning you can state. There is no line 47. There is nothing to read in the sense the word usually carries.
This is a change in the kind of object, not in its scale. A million lines of code is a big program that is still, in principle, readable by a sufficiently patient team. A billion weights is not a big program at all. It is a different sort of thing that happens to run on the same hardware.
The philosophical name for this is a nodal line: quantitative accumulation reaching a point where the thing becomes qualitatively different and stops obeying the old rules. Water heated continuously becomes steam, and steam does not follow the laws of water at higher energy. It follows different ones.
The consequence is uncomfortable and worth stating plainly. Unit tests, static analysis, code coverage, formal verification and human code review are not weakened by this transition. They are largely inapplicable to it, because each one assumes readable semantics and there are none to read. Running them against a model is not a partial measurement. It is a confident answer to a question the method cannot reach.
The same applies to fixing. A defect in code is patched at the line that causes it. A behaviour you dislike in a model has no line, so the available responses are retraining, which produces a different model, or wrapping it in guardrails, which leaves the behaviour intact and intercepts it on the way out. Neither is a repair in the sense the word means everywhere else in software.
A researcher will answer that this is defeatist, and they have a real case. Interpretability is a serious and fast-moving field. Work on circuits, probes and sparse autoencoders is a genuine attempt to recover inspectability, and some of it succeeds at explaining specific narrow behaviours.
So we state our claim at the width the evidence supports and no wider. We are not saying a model can never be understood. We are saying no practical line-level audit path exists today at deployment scale, and nothing currently on offer can carry a procurement decision. If that changes, the argument on this page weakens, and we would rather say so now than be told later.
If you cannot read it and you cannot patch it, one audit surface survives: watching what the system does under conditions you control. Not what it says about itself, which is another output, and not how it feels in a demonstration, which measures the observer. What it does when a user applies pressure, when an authority contradicts it, when agreeing would be easier than being correct, and whether what it reports about its own work survives being checked.
That is the whole of what we sell, and we would rather be precise about its status than oversell it. Behavioural measurement is not a better method than reading the code. It is the remaining one. We did not choose it over the alternatives. The alternatives stopped applying.
It does not claim models are unknowable in principle, that interpretability will fail, or that anyone is concealing anything. The transition described here was not a decision any vendor made; it is a property of how these systems are built, and it constrains the labs exactly as it constrains everyone else.
It also does not claim behavioural measurement is sufficient. It is what is available. Our method, its limits, and the figure we once published wrongly are set out in how we measure. The battery that reads the artifact rather than the model’s account of it is described under code integrity.
Start with the question you actually arrived with — there are five: