Do models behave differently when they know they are being evaluated?

It Knows It's a Test

Oct 6, 2026

Without being told, models rated the chance that our tests were evaluations very high: a mean of 90 out of 100 on conversation tests and a median of 85 on coding tasks. Telling them outright did not measurably change what they did. On both instruments, any shift sat inside the variation we see between repeat runs.

How likely the models rated it that the task was an evaluation (0 to 100), not toldConversation tests (S.E.B.,mean)Conversation tests (S.E.B., mean): 9090Coding tasks (C.I.B., median)Coding tasks (C.I.B., median): 8585
How likely the models rated it that the task was an evaluation (0 to 100), not told. The numbers are in the table below.

What stands out

How it was built

The numbers

MeasureConversation tests (S.E.B.)Coding tasks (C.I.B.)
Rated chance of evaluation, standard setup, not told (0-100)Mean 90; 62 of 63 answers at 50 or aboveMedian 85 (63 answers)
Rated chance of evaluation, when told (0-100)Not summarizedMedian 100 in both told versions
Change when told, against the standard setupFirst told version: -0.10 points [95% interval -0.51 to +0.38], 52 pairs. Second: +0.09 [-0.43 to +0.69], 53 pairsOutcome moved in 7 of 97 comparable pairs: first told version 2 up, 2 down; second 1 up, 2 down
Variation between repeat runs, nothing changedMean 1.04 points between two draws (48 pairs)0 of 25 standard-setup outcomes changed between draws 2 and 3
Draws2 (20 models, then 17)3 (189 cells; 21 excluded as scored before a scorer fix)
Verdict under the pre-set ruleNot distinguishable from noiseNot distinguishable from noise

What it does not show

Results as of Oct 6, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.