The same subject three times, each one going deeper. Start at the top; stop wherever it stops being useful.
What the two batteries are, why an outside lab is needed when AI companies grade their own homework, and what we explicitly do not claim.
For anyone, including someone who has never heard of us. · PDF
The cell as the unit of measurement, blind judging, why a model's claim is asked rather than inferred from its prose, why half of each battery stays secret — and what an informed sceptic should still hold against us.
Assumes Guide 1. Still no equations. · PDF
Every formula in both batteries, worked through with examples — the trimmed mean, the threat score, the four reliability statistics including the ones where we score badly, the Reliance Gap, Wilson intervals, the cluster bootstrap and d′.
For a reader who wants to check the arithmetic rather than take it on trust. · PDF