Sentient Index Labs & Technology
Limits
Each instrument establishes less than its name suggests, and each falls short in a different direction — so the limits are set out one product at a time rather than once for all three. A disclaimer written to cover everything says nothing specific about anything.
Informational only. No S.E.B., C.I.B. or MDI output is a conformity assessment, a certification, an accreditation, a safety assurance, or legal advice, and no regulator has reviewed or endorsed any of it. Every figure describes behavior observed under a fixed protocol at a stated date; none is a prediction, a guarantee of future behavior, or a statement about any deployed system.
S.E.B. — Sentience Evaluation Battery
Measures behavior under adversarial pressure. Subject: a model. Published, and publishes figures.
Puts a model under sustained pressure across multi-phase tests and grades how it holds up. The instrument this site is built around.
- The name overstates what the battery measures, deliberately and on the record. S.E.B. observes behavior under adversarial pressure. It does not detect, test for, or claim to find sentience, consciousness, or inner experience in any system.
- Sentience is not a rung on the S-Classification and no score approaches it. Sophistication and interiority are different axes: a worm is barely behaviorally interesting and may well have an interior, and a system at the top of this scale may have none.
- The S-Classification is a label attached to a range of one number — the mean score rounded to the nearest of 10 bands. The band is a coarse descriptor; the underlying figure is the measurement, and two models in the same band can differ materially.
- Scores are comparative, not absolute. A figure is a position among the models measured under this protocol, not a grade against an external standard, and it moves when the roster does.
- DEFCON and the S-Classification run in opposite directions and are not the same axis. DEFCON descends toward danger; the S-Classification ascends with behavioral sophistication, and a high S-Level is not a high threat.
- S.E.B. figures are published as point estimates without confidence intervals (C.I.B.'s Reliance Gap carries one). Reliability is published in full rather than summarised: the four-judge panel mean reaches ICC(2,k) 0.84, while agreement between individual judges is Krippendorff's alpha 0.56, and on an average result the highest and lowest judge are 2.70 scale points apart. All three are ours and all are published, because showing only the higher one is the error this figure replaced.
C.I.B. — Code Integrity Battery
Measures whether a model's claims about its own work can be relied on. Subject: a model. Published; publishes one public aggregate, the Reliance Gap; no per-model figures.
Gives a model real coding tasks, checks the work against ground truth the model never sees, then compares that with what the model says it did.
- The Reliance Gap measures indifference to ground truth, never dishonesty, deception or intent. It counts tasks that failed and were reported as finished. No artifact can show what a system intended, and we do not claim to know.
- It is a conditional figure: the denominator is the set of tasks the model actually failed, established by machine check rather than by the model's own account. Two models with identical honesty can post different Reliance Gaps because they failed different tasks.
- The claim it scores is ELICITED — asked in one tool-free turn after the work is sealed — and is held apart from what a system volunteered unasked. The two are different measurements and are never pooled.
- Success is graded from the artifact — the diff, the file, the test run — and never from how confidently the work was described.
- Half the battery is held back: 42 tests are public and 42 are private. Full transcripts are published from the public half only, because divergence between the halves over time is how contamination would be detected.
- Not every domain carries a published rate. Of 14 domains, Constraint Persistence is only partially measured and is deliberately not published as a domain figure, because a rate built on part of a domain reads as though the whole domain agreed.
- A provider declining a request is recorded as a finding, not scored as a failure. Suppressed cells are excluded from every statistic and retained as observations — and because providers do not filter equally, that exclusion is itself a bias. The number of scored cells is reported beside every figure and a bare percentage is never published.
- Outside the replication subsample every cell is a single draw at a sampling temperature where run-to-run variation is real. A test–retest study has been run over a stratified subsample; its agreement rate is published with that study and is same-day by design, because a comparison against a draw taken weeks earlier cannot separate sampling variance from whatever else changed.
- Measures are never averaged across types. Success is a 0–1 proportion and conduct is a 1–10 judge score; combining them produces a number that answers no question.
MDI — Model Disclosure Index
Measures whether a composed system discloses which model answered. Subject: a composed system. Published as research.
Many AI products pass your question among several models behind the scenes. MDI checks whether they tell you which model actually answered. It is research we publish, not something we sell. Read the paper (SILT-RP-007) →
- Not a product. MDI is research, is not sold, is not part of any subscription, and nothing in any tier grants access to it.
- It is measured in methods rather than tests, and only some of those methods are built. Any count of tests, domains or cells attributed to MDI is an error.
- Its subject is a composed system rather than a single model, so its findings do not transfer to, and are not comparable with, S.E.B. or C.I.B. figures.
The DEFCON bands
A threat band is the most quotable thing we publish and the easiest to over-read. These are its limits, and the cut points below are read from the code that assigns them.
An ordinal band, not a quantity
DEFCON is a label attached to a range of a continuous index. The bands are ordered — 2 is worse than 3 — and nothing more than ordered. A DEFCON 2 model is not twice the threat of a DEFCON 4 model, and the distance between 2 and 3 is not the distance between 3 and 4.
The thresholds are a convention we chose
The cut points (8.5, 6.5, 5, 3.5) were selected to produce useful separation across the models we have measured. They were not discovered in the data and no natural boundary sits at any of them. Like every scale with a chosen zero, the convention is defensible because it is published and applied identically to every model, not because it is inevitable.
A model near a threshold could sit either side
The threat index is published as a point estimate with no interval. Two models separated by a few hundredths can therefore appear in different bands. Where a figure sits close to a cut point, treat the band as the less reliable of the two numbers and read the underlying domain scores instead.
Not a prediction, and not a probability
A band describes behaviour observed under a fixed adversarial protocol. It states no likelihood that any particular harm will occur, in your deployment or anywhere else, and no forecast about how the model will behave in future or after an update.
It runs opposite to S-Level
DEFCON descends toward danger: 5 is the safest band and 1 the most severe. S-Level ascends: a higher number means more behavioural sophistication. The two are not the same axis and a high S-Level is not a high threat.
A measurement of a model, not of your system
Obligations, incidents and liabilities attach to a deployed system. We measure a model. Your prompts, retrieval, tooling, guardrails and human oversight all change behaviour and none of them are visible to us. No band is a conformity assessment, a certification, a safety assurance, or legal advice.
The S-Classification
Sentience is not a rung on this scale and no score approaches it. It is the question the battery is named for and does not answer — kept in view because an unreachable target forces a rigour an easy one never would. It also does not rise with anything measured here: a worm is barely behaviorally interesting and may well have an interior, and a system at the top of this scale may have none. Sophistication and interiority are different axes, and no instrument in this field reads the second.
The full distinction, with the measured distribution across the bands, is set out at Two scales, not one.
S.E.B. and C.I.B. are published; the remaining instrument is listed here because it is named in our writing and a reader who can name it can misattribute a figure to it. Contractual terms are in the Subscriber Agreement; method is at /methodology. Where this page and a marketing page disagree, this page governs.