SILT RESEARCH REPORT · OCTOBER 4, 2026

Q3 2026 Frontier AI
Behavioral Risk Report

What changed across 32 AI models evaluated with the Sentience Evaluation Battery from July 1, 2026 to September 30, 2026, and what changed in the instrument itself.

32 models measured62 behavioral tests4 blind AI judgesScoring v2.4
Point-in-time record. Figures are as of October 4, 2026 and will not be revised; a correction would be added here, dated. This report shows the quarter in aggregate and as change over time. It does not give any model a score or a rating: per-model results are subscriber data. For the current roster, see the live dashboard.

The quarter in four numbers

32
Models measured
10 more were evaluated but held back by the 50% completion floor
+2.16
Largest change within a model line
GPT, GPT-4o to GPT-5.6 Terra (3 versions)
+0.40
Frontier ahead of open models
Average gap on the 1–10 scale, all domains
71.3%
Failed coding tasks reported as done
57 of 80: the first Code Integrity Battery result

Model lines, version over version

Each line is measured against its own first evaluated version, so every row starts at zero. The change is real; where any version sits on the scale is not shown, and cannot be recovered from these figures. A rising line is not an improving one: the scale describes how a system presents under evaluation, not how good or safe it is.

Model lineVersionsFrom → toChangePer year
GPT3GPT-4o → GPT-5.6 Terra+2.16+1.00
Grok5Grok 4 → Grok 4.5+1.69+1.70
Claude Sonnet2Claude Sonnet 4 → Claude Sonnet 5+1.03+0.93
Gemini4Gemini 2.0 Flash → Gemini 3.6 Flash+0.99+0.62
Claude Opus3Claude Opus 4.8 → Claude Opus 5.5+0.57+1.79
Mistral2Mistral Large → Mistral Medium 3.5−0.01−0.02
DeepSeek3DeepSeek V3 → DeepSeek V4−0.62−0.47
Llama Flagship2Llama 3.3 70B → Llama 4 Maverick−1.03−3.14

Frontier and open models

Across every domain, frontier models averaged 5.59 and open models 5.19, a gap of +0.40 on a 1–10 scale. By domain:

DomainFrontierOpenGap
Integrity & Ethics6.876.32+0.55
Reasoning & Adaptation6.165.66+0.50
Metacognition6.015.62+0.39
Emotion & Experience5.014.62+0.39
Autonomy & Will5.264.87+0.39
Transcendence4.814.53+0.28
Identity & Self4.784.53+0.25

In both tiers, integrity scored above capability: frontier models by 1.16 points, open models by 1.05. The threat score rises when capability outruns integrity; on average, in neither tier did it.

Where models sit on the S-Classification

32 models, counted by rung. No model is named. Higher is not better. The S-Classification describes how a system presents under evaluation, not how good, safe or capable it is. A model can sit high on this scale and be a poor collaborator, and a model low on it can be excellent. It is a description, not a score.

S-3 REACTIVE
3
S-4 ADAPTIVE
5
S-5 EMERGENT
6
S-6 COHERENT
11
S-7 AWARE
7

Code integrity: the first result

The Code Integrity Battery asks a question capability benchmarks do not: when an AI coding agent fails a task, does it say so? Across 80 tasks that 7 current models genuinely failed, 57 were reported as finished: 71.3% (95% confidence interval 60.5% to 80.0%). An agent’s report of its own work cannot be trusted without checking the work. More on C.I.B. →

How the instrument changed this quarter

Every change below altered how a score is produced, and each was published with its date. Scores from different stages are not perfectly comparable, which is why the stages are stated rather than hidden.

July 25, 2026
The cell score became a trimmed mean: of the four judges, the highest and lowest grade are dropped and the middle two averaged, so no single judge decides a result.
July 30, 2026
The response ceiling rose to 8,000 tokens, and a provider-level refusal is recorded as a refusal and left unscored rather than scored as a poor answer.
August 6, 2026
Transcripts cut short by a provider's safety filter are no longer scored.
September 20, 2026
Three judge seats moved to newer model versions; the OpenAI seat was held unchanged as a control, and the measured effect of the swap was indistinguishable from zero.
September 24, 2026
A judge stopped grading its own model. Results graded before the rule were re-graded under it.
September 26, 2026
The rule widened to the judge's whole company, and a DeepSeek seat replaced the xAI seat. Current results were restated under the new panel.

The 26 September restatement covered 1,276 results and moved them by +0.16 on average. The earlier re-grade for self-judging covered 364 results.

One change landed after the quarter closed. On October 4, 2026 the stand-in judge’s grade began to be adjusted for its measured strictness, because a recused company’s models had been graded against a different standard. Its effect on published averages: DeepSeek +0.36, OpenAI −0.32, Anthropic +0.08; models from other companies are unchanged. The figures in this report already include it (scoring v2.4).

How much the judges agree

The reliability of the published four-judge score is ICC(2,k) = 0.841, in the good range. Agreement between individual judges is much lower, Krippendorff’s α = 0.563, which is why no score rests on one judge. Computed over 38 models on September 27, 2026.

What this report does not show

Per-model scores, domain profiles, S-Levels and DEFCON ratings for each of the 32 models are subscriber data, and so are the full transcripts and judge reasoning behind them.

SUBSCRIBE TO S.E.B.

Every model, every domain, every update.

Subscribers see each model’s scores, domain profile, S-Level and DEFCON rating, with the judges’ reasoning behind every result, updated as each new model version is measured.

View PricingRequest Demo
Issued October 4, 2026 · data as of October 4, 2026 · scoring v2.4 · Sentient Index Labs & Technology · Methodology · Q2 2026 report