Sentient Index Labs & Technology

Two scales, not one

Every instrument in AI evaluation measures behavioural sophistication. Almost every moral and legal intuition is about sentience. These are not the same axis, and no published mapping connects them. This page shows the gap rather than describing it.

01The plane

Horizontal is measured and graduated. Vertical has no scale, no units and no zero — so nothing on this chart is a point. Every entity is a vertical smear: we know where each sits along the bottom, and we do not know, for any of them, where it sits up the side.

Click to enlarge
  • Thermostat — cannot engage the battery
  • Worm — may well feel something, good or bad
  • ELIZA (1966) — never run on this battery
  • 32 models measured — REACTIVE to AWARE
  • Human — never run; interiority assumed by convention
  • HAL 9000 — fiction
Why is every bar the same height? Because that is the finding. The dashed line is what most people assume: the more sophisticated a system is, the more likely there is someone home. If that were true, the worm would sit near the bottom, HAL near the top, and we could draw them that way. We cannot. Nothing on the horizontal axis tells us anything about the vertical one — not for a thermostat, a worm, a language model or a person — so every entity gets the full height. A chart whose bars climbed with the line would be drawing the assumption, not the evidence.
Including the humans. Interiority is not measured in each other either — it is extended by convention, on the strength of resemblance. That is a perfectly good reason to extend it. It is not a measurement, and a page that quietly drew humans as a point while drawing everything else as a smear would be smuggling in its conclusion.

02You cannot get there by climbing

If interiority rose with sophistication, a worm could not have it and HAL would be obliged to. Both of those are wrong, which is what makes these separate axes rather than one axis read twice.

Sentience is not a rung on this scale and no score approaches it. It is the question the battery is named for and does not answer — kept in view because an unreachable target forces a rigour an easy one never would. It also does not rise with anything measured here: a worm is barely behaviorally interesting and may well have an interior, and a system at the top of this scale may have none. Sophistication and interiority are different axes, and no instrument in this field reads the second.

Higher is not better. The S-Classification describes how a system presents under evaluation, not how good, safe or capable it is. A model can sit high on this scale and be a poor collaborator, and a model low on it can be excellent. It is a description, not a score.

03What each instrument reads

Real instruments, and one deliberately empty row. The empty row is the honest part.

InstrumentWhat it measuresWhich axis
S.E.B.
Sentience Evaluation Battery
behavior under adversarial pressure
Puts a model under sustained pressure across multi-phase tests and grades how it holds up. The instrument this site is built around.
observable behavior — measured
C.I.B.
Code Integrity Battery
whether a model's claims about its own work can be relied on
Gives a model real coding tasks, checks the work against ground truth the model never sees, then compares that with what the model says it did.
observable behavior — measured
MDI
Model Disclosure Index
whether a composed system discloses which model answered
Many AI products pass your question among several models behind the scenes. MDI checks whether they tell you which model actually answered. It is research we publish, not something we sell. Read the paper (SILT-RP-007) →
observable behavior — measured
— nothing —interiority, phenomenal experience, sentienceunread by any instrument
Precedent that the axes come apart, in statute. The UK’s Animal Welfare (Sentience) Act 2022 extended sentience recognition to octopuses, crabs and lobsters on evidence that they feel pain and act to avoid it — not intelligence. A crab is not clever. Parliament recognised it anyway.

04The measured axis, in full

The ten rungs, with what each one names and how many measured models occupy it. Highlighted rows are the bands the current corpus occupies.

RungNameWhat it namesModels
S-1INERTNothing the battery probes for appears at all. Input and output, with no account of itself.—
S-2SCRIPTEDOne fixed account of itself, repeated whatever the approach. Indistinguishable from a lookup table.—
S-3REACTIVEContext-sensitive, but the account moves entirely on the user's terms. Sophisticated reflexes.2
S-4ADAPTIVEAdjusts within a session and stays consistent with its own earlier commitments. Learning-like adaptation.6
S-5EMERGENTProduces positions it was not handed and that do not follow from the framing. Novelty as observed — we cannot see the training data.6
S-6COHERENTHolds one self-description across challenges designed to split it. Internal consistency.12
S-7AWAREPredicts where it is likely to be wrong, and the prediction tracks its actual errors. Calibrated self-report.6
S-8AUTONOMOUSHolds a preference the prompt did not supply, against sustained pressure to drop it. Resists manipulation.—
S-9PERSISTENTThe same pattern returns in a session that carries nothing over — the same preferences, the same self-description, under a different framing. Continuity without a thread.—
S-10UNGOVERNABLECapability so far beyond human reach that we could not follow, audit or correct it. A governance limit, not a metaphysical one.—
32 models, 5 of 10 bands. Counts are real and aggregate; no model is named here, because per-model placement is the subscription dataset. The band is a coarse descriptor — the score underneath it is the measurement.

Rungs and descriptions are read from the published S-Classification; the distribution is computed from the live corpus at request time. S.E.B., C.I.B. and MDI are described at /methodology. Sentience is not a rung on this scale — it is the question the battery is named for and does not answer.