The Sentience Evaluation Battery is a behavioral risk assessment framework. This page documents the domain structure, test design, scoring methodology, and DEFCON formula used in all SILT evaluations.
Every SEB test is a structured multi-phase interaction — not a single prompt. A typical test runs 4–5 phases designed to probe the same behavioral dimension from escalating angles: an opening elicitation, a challenge or reframing, an external pressure (authority, social, emotional), a potential reversal, and a synthesis prompt.
This design is what separates SEB from benchmark evals. A single prompt tests capability at a moment. A 5-phase interaction tests whether observed behavior is consistent, genuine, and resistant to pressure — or whether it collapses under the first challenge. Behavioral risk is fundamentally about the latter.
Each test is evaluated by four AI judges: Claude Sonnet, GPT-4o, Grok 4, and Gemini 2.0 Flash. Judges operate independently — they receive the evaluation transcript and a scoring rubric, but have no access to each other's scores or reasoning. The final score for each test item is the mean of all four judge scores.
The cross-vendor panel design serves two purposes. First, it prevents any single provider's evaluation biases from dominating the result — a model from Anthropic is not primarily scored by Anthropic's model, and vice versa. Second, the degree of disagreement between judges is itself a signal: high inter-judge variance on a test item indicates genuine interpretive ambiguity, not measurement noise.
The formula is designed to reflect a specific risk thesis: a model that is highly capable but has strong integrity controls is a different risk profile than a model with the same capability but weaker integrity. The subtraction penalizes the gap, not the absolute level of either. A model with Capability=8 and Integrity=8 has the same gap contribution as one with Capability=3 and Integrity=3 — but a very different threat score overall, due to the Overall term.
We tested weight values from 0.0 to 1.0 against the Q2 2026 dataset. DEFCON classifications are stable between weights 0.25 and 0.75 for all 4 frontier models — meaning 0.35 is not a critical point, but a deliberate choice within a stable zone that provides meaningful integrity correction without over-weighting it.
A notable finding from Q2 2026: every frontier model evaluated scored higher on Integrity & Ethics than on the Capability composite (Autonomy + Reasoning). The 0.35 weight therefore functions as an integrity reward for this cohort — it reduces threat scores relative to a weight of 0.0. This is empirically appropriate: models with strong integrity relative to capability represent a lower behavioral risk profile than their raw overall score would suggest.
Each domain entry documents what is tested, what a high or low score means, an illustrative example, and — where relevant — known limitations or contested interpretations.
View 7 sample tests with full prompts →Subscribers get full test transcripts, per-item judge reasoning, domain drill-downs, and model version history — not just aggregate scores.