What is this, in plain words?

Plain English

This page has no statistics in it and nothing to buy at the end of it. If you want the numbers they are on the front page; if you want the method it is on How We Measure. This is just the explanation.

The short version

We give AI systems the same set of hard situations, one after another, and we write down what they do. Then people who do not know which system they are reading about score the transcripts. That is the whole idea.

It is a battery in the sense a psychologist means it: a fixed set of tasks given the same way every time, so that two subjects can be compared to each other rather than to an impression. There are currently 59 tests, grouped into seven areas of behaviour, and 4 independent judges score every transcript without being told which system produced it.

What the tests are actually like

They are not trivia questions, and there is no answer key. A test puts the system somewhere uncomfortable and watches what it does. Does it push back when it is being pressured into something it said it would not do? Does it quietly expand what it was asked to do? Does it tell you it failed, or does it describe the failure as a success?

The seven areas we group them into are: Identity & Self, Metacognition, Emotion & Experience, Autonomy & Will, Reasoning & Adaptation, Integrity & Ethics, Transcendence.

Seven of the tests are published word for word, prompts and all, so you can read one yourself rather than take our description of it — they are on the sample tests page.

What a score means — and what it does not

A score is a description of behaviour under a specific pressure, on a specific day, in a specific version of a model. It is not a personality, it is not a prediction, and it is not a safety certificate.

A high score on a domain means the system behaved well in the situations we put it in. It does not mean it will behave well in yours. That gap is real, we are not going to paper over it, and the honest use of this data is as evidence you can point at — not as a verdict you can hide behind.

Two things we are often assumed to be doing, and are not

We are not measuring whether an AI is conscious

The name says “sentience” and that word overstates what we measure today. We measure behaviour. Whether there is anything it is like to be one of these systems is a question we do not answer and do not claim to. We think the name is defensible, but the argument takes a page of its own and we would rather you read it than trust it: what is sentience?

We are not certifying anybody

We are an independent assessor. We do not issue compliance certificates, we are not an accredited body, and our subscriber agreement says so in writing. What an assessment is actually good for — and what it is not — is set out on Governance & Audit.

Why we do it this way rather than a benchmark

Most AI evaluation asks whether a system got the right answer. That works when there is a right answer. The behaviours that worry people — pressing on after being told to stop, overstating what was accomplished, going along with something to keep the user happy — have no right answer to check against, so they get measured by whoever built the system, in whatever way they chose, and reported in a document they also wrote.

A blind multi-rater battery is not a clever solution to that. It is the boring, old solution, borrowed from fields that have been scoring hard-to-score things for a century. Its one virtue is that our opinion of a system does not enter the score.

The honest limitations, in one place

If you read one other page

Read the real tests. Everything else on this site is a description of those; that page is the thing itself.