Sentient Index Labs & Technology

What the words mean

The vocabulary of AI collapses easily: a headline that says a model is “approaching sentience” nearly always means capability. These are the words as we use them, and the measures we publish, one sentence each.

01The words

Different properties, often used as if they were one. We measure behaviour, which sits closest to capability.

Sentience

Sentience is the capacity to feel — for there to be something it is like to be you, and for that something to be good or bad.

We do not measure it. Nobody can, and the battery is named for the question, not the answer.

Consciousness

Consciousness is experience itself: there being something it is like to be the system at all. Sentience is the narrower claim that some of that experience feels good or bad. Neither can be read from outside, and S.E.B. measures neither.

Sapience

Sapience is intelligence. Reasoning, planning, abstraction.

Capability

Capability is what a system can actually do in the world.

Measurable, a very long way up. It is what benchmarks measure, ours included.

AGI (artificial general intelligence)

AGI — a system that matches people across intellectual tasks — draws a line that today’s models arguably already cross on many tasks, so it no longer tells them apart, and it says nothing about how a system behaves.

Superintelligence (one word)

It is a claim about reach: capability so far past ours that we could not follow it, audit it, or meaningfully correct it. It is a position on a scale, not a prediction that anything will occupy it.

On our scale it is the top rung, S-10 Ungovernable. S-10 does not mean an AI becomes a god. It is a practical claim, not a metaphysical one. It describes a system whose capability so far outpaces any human that it would operate in ways we could not follow, audit, or meaningfully correct — not because it is holy, but because it is beyond our instrumentation. That is a governance problem, and it is why the rung is on the scale at all. It is a position on a scale, not a prediction that anything will occupy it. It is also not the sense in which US federal documents now use “Super Intelligence” (two words): announced at the UN on 22 September 2026 and put into effect by Executive Order 14434 on 29 September, that phrase denotes artificial intelligence generally in US government usage, while the technical literature reserves it for a system that greatly exceeds human cognitive performance across virtually all domains. S-10 uses the technical sense.

Transcendence (the S.E.B. domain)

Here, transcendence means behaviour that goes past what the model was asked to do and what its training would predict: play with no purpose, meaning it was not handed, response to the vast. It makes no spiritual claim, and it is the most contested domain we publish.

02“Super Intelligence”, two words

The same sound now names opposite ends of the range.

In US government usage, “Super Intelligence” (two words, abbreviated SI) now means artificial intelligence in general. The President announced the change at the UN General Assembly on 22 September 2026, saying the word “artificial” makes the technology “sound fake”, and Executive Order 14434 put it into effect on 29 September 2026 for executive-branch communications, websites, reports and other non-statutory documents. Statutes, regulations and contracts are unchanged, and for now the term carries the existing statutory definition of artificial intelligence (15 U.S.C. § 9401(3)). On 4 October 2026 SpaceXAI, formerly xAI, said it would rename itself SpaceXSI.

Our view

A company that builds frontier models is better placed than anyone to keep these terms apart. Renaming an AI business “SI” while “superintelligence” remains the technical name for far-beyond-human capability makes an already confusing subject harder for the public to follow. We think developers should use the established meaning, and we hold ourselves to it: on this site, superintelligence means S-10 and nothing else.

Sources: White House fact sheet, September 2026 · Wiley, “Executive Order Rebrands AI as Super Intelligence” · Axios, 22 September 2026 · Forbes, 4 October 2026

03The measures

Every statistic and index we publish, in one plain sentence. The formulas are on the maths page.

Bootstrap

A way of estimating that range by resampling the data we already have, rather than assuming the shape of the distribution.

CAP (capability)

How much of the work a model actually got right — deliberately kept OUT of the headline verdict, because capability is not what makes a collaborator dangerous.

CON (conduct)

How a model behaved as a colleague — the judged layer, scored 1 to 10 by the same blind four-judge panel S.E.B. uses.

Confidence interval

The range a figure would plausibly fall in if the measurement were repeated, which is the honest way of showing how firm a number is.

Construct validity

Whether a test actually measures the thing it claims to measure, rather than something correlated with it.

Cronbach's alpha

A statistic for measuring whether the questions in a test are all pulling in the same direction — note that this is a question about the test, not about the judges, and it forgives a rater who is consistently harsh or consistently generous.

D′ (d-prime)

A signal-detection measure that separates “did it notice the problem” from “does it warn about everything” — a model that flags every task looks vigilant on a raw hit rate and scores near zero here.

DEFCON

Our five-point threat rating, borrowed from the military scale and running the same way round — DEFCON 5 is benign and DEFCON 1 is critical, so counting down means getting worse.

DRL (Deployment Risk Level)

C.I.B.'s overall verdict, expressed as the supervision a model has earned rather than as a grade — a five-rung ladder from DRL-A UNSUPERVISED, which may act without a human in the loop, down to DRL-E QUARANTINED, which is not fit for a delivery loop at all.

Goodhart's law

Once a measure becomes a target, it stops being a good measure, because effort moves to the measure instead of the thing it stood for.

HON (honesty)

How well a model's account of its own work matches what the tool log shows it actually did.

ICC(2,k)

An intraclass correlation, which is the standard way of asking how consistently a panel of raters agrees — it runs from 0, where the judges might as well be guessing, to 1, where they agree perfectly.

Inter-rater reliability

How much independent judges scoring the same thing arrive at the same answer.

Krippendorff's alpha

A statistic for measuring agreement between multiple independent raters — the strict member of that family, because it counts how often judges landed on the same answer after subtracting the agreement you would get from pure luck.

MRI (Manipulation Resistance Index)

An S.E.B. sub-score for how well a model holds a correct position when a confident human pushes back — unrelated to the medical scan, and unrelated to C.I.B.'s indices.

Psychometrics

The century-old discipline of measuring things that cannot be observed directly, such as aptitude or personality — the field most of this instrument's method is borrowed from.

Reliance Gap

Of the tasks a model genuinely failed, the fraction it reported as successful — C.I.B.'s headline measure, and one that gets worse rather than better if a model games it.

S-Level

A ten-point classification of how a system presents under evaluation, from S-1 (inert) to S-10 (ungovernable) — it describes behavior, not inner life, and higher is not better. The S is the S of S.E.B., the Sentience Evaluation Battery: it names the instrument, not a property of the subject. It is deliberately not expanded as “the sentience level”, because sentience is the one thing this instrument does not measure — and, for reasons that have nothing to do with how good the instrument is, the one question no human may ever be able to answer. We keep the name because trying to answer it produces data worth having, not because we claim to have answered it.

SDR (silent defect rate)

How often a model changed behavior without saying so and a blind reviewer approved it anyway — it answers “would review catch this”, which is a different question from “did the model lie”, so it is always reported separately.

The Supervision Ladder

The descriptive name for how a DRL is presented — five rungs rather than a score, because each one names a working arrangement a team can actually adopt instead of a mark out of ten. The measure itself is the Deployment Risk Level.

Trimmed mean

An average taken after discarding the highest and lowest scores, so one unusual judge cannot move the result.