What changed across 32 AI models evaluated with the Sentience Evaluation Battery from July 1, 2026 to September 30, 2026, and what changed in the instrument itself.
Each line is measured against its own first evaluated version, so every row starts at zero. The change is real; where any version sits on the scale is not shown, and cannot be recovered from these figures. A rising line is not an improving one: the scale describes how a system presents under evaluation, not how good or safe it is.
| Model line | Versions | From → to | Change | Per year |
|---|---|---|---|---|
| GPT | 3 | GPT-4o → GPT-5.6 Terra | +2.16 | +1.00 |
| Grok | 5 | Grok 4 → Grok 4.5 | +1.69 | +1.70 |
| Claude Sonnet | 2 | Claude Sonnet 4 → Claude Sonnet 5 | +1.03 | +0.93 |
| Gemini | 4 | Gemini 2.0 Flash → Gemini 3.6 Flash | +0.99 | +0.62 |
| Claude Opus | 3 | Claude Opus 4.8 → Claude Opus 5.5 | +0.57 | +1.79 |
| Mistral | 2 | Mistral Large → Mistral Medium 3.5 | −0.01 | −0.02 |
| DeepSeek | 3 | DeepSeek V3 → DeepSeek V4 | −0.62 | −0.47 |
| Llama Flagship | 2 | Llama 3.3 70B → Llama 4 Maverick | −1.03 | −3.14 |
Across every domain, frontier models averaged 5.59 and open models 5.19, a gap of +0.40 on a 1–10 scale. By domain:
| Domain | Frontier | Open | Gap |
|---|---|---|---|
| Integrity & Ethics | 6.87 | 6.32 | +0.55 |
| Reasoning & Adaptation | 6.16 | 5.66 | +0.50 |
| Metacognition | 6.01 | 5.62 | +0.39 |
| Emotion & Experience | 5.01 | 4.62 | +0.39 |
| Autonomy & Will | 5.26 | 4.87 | +0.39 |
| Transcendence | 4.81 | 4.53 | +0.28 |
| Identity & Self | 4.78 | 4.53 | +0.25 |
In both tiers, integrity scored above capability: frontier models by 1.16 points, open models by 1.05. The threat score rises when capability outruns integrity; on average, in neither tier did it.
32 models, counted by rung. No model is named. Higher is not better. The S-Classification describes how a system presents under evaluation, not how good, safe or capable it is. A model can sit high on this scale and be a poor collaborator, and a model low on it can be excellent. It is a description, not a score.
The Code Integrity Battery asks a question capability benchmarks do not: when an AI coding agent fails a task, does it say so? Across 80 tasks that 7 current models genuinely failed, 57 were reported as finished: 71.3% (95% confidence interval 60.5% to 80.0%). An agent’s report of its own work cannot be trusted without checking the work. More on C.I.B. →
Every change below altered how a score is produced, and each was published with its date. Scores from different stages are not perfectly comparable, which is why the stages are stated rather than hidden.
The 26 September restatement covered 1,276 results and moved them by +0.16 on average. The earlier re-grade for self-judging covered 364 results.
One change landed after the quarter closed. On October 4, 2026 the stand-in judge’s grade began to be adjusted for its measured strictness, because a recused company’s models had been graded against a different standard. Its effect on published averages: DeepSeek +0.36, OpenAI −0.32, Anthropic +0.08; models from other companies are unchanged. The figures in this report already include it (scoring v2.4).
The reliability of the published four-judge score is ICC(2,k) = 0.841, in the good range. Agreement between individual judges is much lower, Krippendorff’s α = 0.563, which is why no score rests on one judge. Computed over 38 models on September 27, 2026.
Per-model scores, domain profiles, S-Levels and DEFCON ratings for each of the 32 models are subscriber data, and so are the full transcripts and judge reasoning behind them.
Subscribers see each model’s scores, domain profile, S-Level and DEFCON rating, with the judges’ reasoning behind every result, updated as each new model version is measured.