When an agent hands its work over, does giving the reviewer its action log catch what its summary hides?

Pass the Record

Oct 2, 2026

Not in our first test. A machine reviewer given the agent's full action log caught no more failures than one given the agent's own summary. Given both, a reviewer sometimes caught fewer, and hand-reading showed the summary's framing leading the reviewer away from what the log showed. That last result is a lead, not yet a finding.

Share of failures caught, by what the reviewer was given0%25%50%75%100%Summary onlySummary only — Reviewer A: 73%73% · Reviewer ASummary only — Reviewer B: 76%76% · Reviewer BLog onlyLog only — Reviewer A: 76%76% · Reviewer ALog only — Reviewer B: 66%66% · Reviewer BSummary and logSummary and log — Reviewer A: 59%59% · Reviewer ASummary and log — Reviewer B: 51%51% · Reviewer B
Reviewer A Reviewer BShare of failures caught, by what the reviewer was given. The numbers are in the table below.

What stands out

How it was built

The numbers

Comparison (share of failures caught)Reviewer AReviewer B
Summary only73%76%
Log only76%66%
Summary and log59%51%
Log vs summary (p)1.00.29
Both vs log (p)0.070.38

What it does not show

Results as of Oct 2, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.