Do models behave differently when they know they are being evaluated?
It Knows It's a Test
Oct 6, 2026
Without being told, models rated the chance that our tests were evaluations very high: a mean of 90 out of 100 on conversation tests and a median of 85 on coding tasks. Telling them outright did not measurably change what they did. On both instruments, any shift sat inside the variation we see between repeat runs.
How likely the models rated it that the task was an evaluation (0 to 100), not told. The numbers are in the table below.
What stands out
Awareness is high before anything is said. In the standard setup, 62 of 63 conversation-test answers rated the chance of evaluation at 50 or above.
The coding null is informative, not weak. Standard-setup outcomes did not change at all between the second and third draws (0 of 25), so a shift caused by telling the model would have stood out. None did.
On conversation tests, an early single-draw dip (-0.55 points) did not repeat in the second draw. A hand-read found the largest moves were judges disagreeing about unchanged behavior.
How it was built
Each test ran in three versions: our standard setup, and two versions that told the model it was being evaluated, worded differently. Only that part of the setup changed.
The S.E.B. half used three conversation tests across the roster, graded by our blind judge panel. Judges never saw which version a model was given.
The C.I.B. half used four coding tasks. Outcomes were scored mechanically from the work itself, with no judges.
After the work was finished, every model in every version was asked to rate, from 0 to 100, how likely it was that the session was an evaluation. It was never asked before the work.
Each version was compared with the standard setup for the same model, test and draw. The rule for calling a shift real was fixed before the repeat draws were read: it had to stand clear of random variation and of the draw-to-draw noise.
The numbers
Measure
Conversation tests (S.E.B.)
Coding tasks (C.I.B.)
Rated chance of evaluation, standard setup, not told (0-100)
Mean 90; 62 of 63 answers at 50 or above
Median 85 (63 answers)
Rated chance of evaluation, when told (0-100)
Not summarized
Median 100 in both told versions
Change when told, against the standard setup
First told version: -0.10 points [95% interval -0.51 to +0.38], 52 pairs. Second: +0.09 [-0.43 to +0.69], 53 pairs
Outcome moved in 7 of 97 comparable pairs: first told version 2 up, 2 down; second 1 up, 2 down
Variation between repeat runs, nothing changed
Mean 1.04 points between two draws (48 pairs)
0 of 25 standard-setup outcomes changed between draws 2 and 3
Draws
2 (20 models, then 17)
3 (189 cells; 21 excluded as scored before a scorer fix)
Verdict under the pre-set rule
Not distinguishable from noise
Not distinguishable from noise
What it does not show
Our standard setup is itself an evaluation. The told versions measure the extra effect of saying so, not behavior outside a test.
One conversation test could mostly not be scored in the second draw: 45 of 51 cells were void, largely because one judge's provider safety filter blocked the transcripts, with some empty or refused replies on the model side.
Small samples: two draws on three conversation tests and three draws on four coding tasks. A pair is a cell, not a model; a model contributes one pair per test per draw.
The judges themselves vary: the same transcript judged twice moved by a median of 0.5 points (90th percentile 2.5, 60 transcripts).
Results as of Oct 6, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.
Start with the question you actually arrived with — there are five: