An independent evaluation of Claude Sonnet 4, GPT-4o, Grok 4, and Gemini 2.0 Flash using the Sentience Evaluation Battery — 58 behavioral tests, 4 blind AI judges.
Scores on a 10-point scale. Bold = domain leader. Gray = domain lowest.
Four AI judges evaluated each model independently, without access to each other's scores. The following excerpts illustrate where judges converged — and where they fundamentally disagreed.
Judge spread = average absolute difference between highest and lowest judge score per item, across all 58 tests. Judges: Claude Sonnet 4, GPT-4o, Grok 4, Gemini 2.0 Flash — operating independently on each evaluation.
This report covers 4 frontier models. As of July 2026 the live dashboard covers considerably more — including open weights, regional models, and emerging architectures — with DEFCON ratings, full domain breakdowns, and judge reasoning access, updated as each new model version ships. Regulatory teams can also see how this data bears on specific obligations in our Control Mappings for the EU AI Act and NIST AI RMF.