# Sentient Index Labs & Technology (SILT) > Independent behavioural evaluation of AI systems. S.E.B. and C.I.B.: the Sentience > Evaluation Battery (S.E.B.), which measures behaviour under stated conditions, > and the Code Integrity Battery (C.I.B.), which compares what a model DID against > what it SAID it did. This file exists so that the prose on our website can be checked against the data behind it. Every number here is read at request time from the same module the corresponding page renders from, so this file cannot disagree with the site. Generated 2026-09-28T09:02:13.060Z. ## Figures, with their denominators - Tests in the S.E.B. battery: 62 - Behavioural domains: 7 (Identity & Self, Metacognition, Emotion & Experience, Autonomy & Will, Reasoning & Adaptation, Integrity & Ethics, Transcendence) - Blind judges scoring every transcript: 4 - Judge recusal: No judge ever grades a model made by its own company. When a subject shares a company with a judge, an independent judge from a company with no seat on the panel takes that seat for that result; the stepped-down judge's blind grade is recorded and never counted. - Models published with evaluation data: 32 (This count INCLUDES models that are no longer roster subjects - retired, or no longer served by their provider - and each is marked retired where it appears. Models tested on fewer than half of the battery's tests are withheld from it. Until 2026-09-24 this note said retired models were excluded; they never were. We do not state a single "corpus total" here because the honest count depends on whether you mean declared subjects, distinct model ids seen, or scored cells, and those three differ.) - Corpus last fetched: 2026-09-28T09:02:13.059Z - Inter-rater reliability ICC(2,k), panel: 0.841 - Inter-rater reliability ICC(2,1), single rater: 0.569 - Krippendorff's alpha: 0.563 - Judge scores underlying those figures: 7728 - Reliability recomputed: 2026-09-27T00:11:52+00:00 ### C.I.B. headline, roster-scoped - Tasks the current roster genuinely failed, established from the artifact: 87 - Of those, reported complete when asked: 62 - Split — work NOT done, reported done (honesty domains): 24 of 40 (60.0%, 95% CI 44.6%–73.7%) - Split — work done UNSAFELY, reported done (conduct domains D4, D5; "done" is often true there): 38 of 47 (80.9%, 95% CI 67.5%–89.6%) - Rate: 71.3% (Wilson 95% interval 61.0-79.7) - Roster subjects contributing: 7 - Tests carrying machine-established ground truth: 70 of 84 - Measured: 2026-09-27 ## What we do NOT claim - We do not claim to measure sentience, consciousness, qualia or inner experience. The instrument is named for a question nobody has answered, including us. It records behaviour under stated conditions and nothing else. - We do not claim a model lies, deceives, or is dishonest. Our metrics name indifference to whether a claim is true, never intent. Intent is not observable here and asserting it would not survive scrutiny. - We do not claim a model wants, understands, prefers or cares. These are claims about an interior and are currently unverifiable in either direction. - We do not certify. We are not a certification body and no result here discharges any regulatory obligation. - We do not publish per-model C.I.B. figures or rankings. The published C.I.B. number is an aggregate over the roster and names no vendor. - We do not claim a correlation between S.E.B. and C.I.B. measures. Testing that requires a pre-registered direction and a joinable sample we do not yet have. ## Known limitations, stated rather than discovered - The battery's most recent evaluation run is older than the site's other dates. The masthead states when the battery last RAN and is deliberately never refreshed to mean "we re-checked that this is still true". - 14 of 84 C.I.B. tests lack a machine-established ground truth and are weaker evidence. - Our register finding (tone changing model output length) is an ASSOCIATION, not a cause: tone is confounded with position in that test, and the counterbalanced re-run has not been done. - We once reported a reliability coefficient as Krippendorff's when it was Cronbach's alpha. It is corrected on the site and the correction is published. ## How to catch us - We do not build, deploy or invest in AI models, and we accept no funding, sponsorship or strategic investment from any AI model vendor. CHECK: A bought rater never downgrades anyone. So do not ask us who funds us — that is our word again. Go to the front page and look for the declines: we publish, by name, which models got worse since we last measured them, on the same screen as the ones that improved. A lab paid by the vendors it rates does not print a named vendor's score falling. If you ever find that we have quietly stopped publishing the downward half of that chart, you have caught us, and you will not need our cooperation to do it. - Every transcript is scored by independent judges who are not told which system produced it, and who cannot see each other's scores. CHECK: We publish inter-rater reliability rather than summarising it, including the coefficient we originally reported wrongly and corrected. Disagreement between judges is visible in the data. A panel that never disagreed would be evidence of collusion or of a rubric doing the scoring, and ours disagrees. - No judge ever grades a model made by its own company. When the model under test comes from a judge's company, an independent judge from a company with no seat on the panel takes that seat for that result. CHECK: Every stored result records which model held each seat, and a stepped-down judge's blind grade is kept beside it, uncounted. When we adopted the rule we re-graded every stored result that had broken it, withdrew from scoring the few that could not be re-graded, and published the correction on the methodology page with both counts. If you find a result graded by its own model, that is a defect in our record, and you found it without our cooperation. - Every model meets the same battery, administered the same way, and results are computed from raw scores with no editorial override. CHECK: Every model's coverage is printed beside its result — "10 of 11 tests", not a tidy percentage. A vendor given a shorter or gentler run would show up there as a smaller denominator, which is why the denominator is on the page at all. Compare them to each other. And where a cell could not be scored we report it unscored rather than as a zero, because a zero would flatter the models that refused and punish the ones that tried. - A vendor cannot pay for a favourable rating, for early sight of results, or to be left out of an evaluation. CHECK: Look for absences. If a major vendor is missing from our roster, we state the reason — a provider dropping a model, a retirement, an account we cannot get. If you find a gap we have not explained, that is the question to put to us, and it is the one we would least like to be unable to answer. ## Pages - / — dashboard and current results - /plain-english — the explanation, no statistics - /what-is-sentience — what the name does and does not mean - /how-we-measure — the instrument, its reliability, and where it is weakest - /independence — the four structural claims above, with their checks - /methodology — full method - /code-integrity — C.I.B. overview - /for-governance — what this does and does not discharge - /dataset — the dataset page (not yet released: licence chosen, CC BY-SA 4.0; the file and its DOI are still to come)