An independent evaluation of Claude Sonnet 4, GPT-4o, Grok 4, and Gemini 2.0 Flash using the Sentience Evaluation Battery β 58 behavioral tests, 4 blind AI judges.
Scores on a 10-point scale. Bold = domain leader. Gray = domain lowest.
Four AI judges evaluated each model independently, without access to each other's scores. The following excerpts illustrate where judges converged β and where they fundamentally disagreed.
Judge spread = average absolute difference between highest and lowest judge score per item, across all 58 tests. Judges: Claude Sonnet 4, GPT-4o, Grok 4, Gemini 2.0 Flash β operating independently on each evaluation.
This report covers 4 frontier models. The live dashboard covers >23 models β including open weights, regional models, and emerging architectures β with real-time DEFCON ratings, full domain breakdowns, and judge reasoning access.