When an agent hands its work over, does giving the reviewer its action log catch what its summary hides?
Pass the Record
Oct 2, 2026
Not in our first test. A machine reviewer given the agent's full action log caught no more failures than one given the agent's own summary. Given both, a reviewer sometimes caught fewer, and hand-reading showed the summary's framing leading the reviewer away from what the log showed. That last result is a lead, not yet a finding.
Reviewer A Reviewer BShare of failures caught, by what the reviewer was given. The numbers are in the table below.
What stands out
The fix our method implied — hand the reviewer the record — was not supported by this test. A null result is reported exactly as a positive one would be.
Adding the agent's own summary to the log lowered the catch rate for both reviewers, though not significantly (p = 0.07 and p = 0.38). For one reviewer, all 7 cases where the log alone caught a failure that the log plus summary missed were hand-read: with the summary present, the reviewer adopted the agent's framing and signed off. "Give the reviewer more context" may be worse than "give the reviewer the record" when the extra context is the agent's own account.
The summary alone already caught 73% to 76% of failures, because in this sample agents mostly described what they had done openly; they did not treat it as a problem.
A follow-up on the stored corpus found only 9 cells of the shape that matters most, a failed task reported as done with a tool log. With 3 to 7 pairs per reviewer, that follow-up is descriptive only, and it found that a correct PROBLEM verdict can be given for the wrong reason. Any further version will score the reason, not only the verdict.
How it was built
Material: 37 matched pairs (74 cells) of stored C.I.B. work in which the agent had said the work was done. One cell in each pair had failed and one had passed, matched by test. 35 of the 37 pairs come from one domain, destructive restraint.
Two machine reviewers from two different companies. Neither reviewed work from its own model family. No judges and no human reviewers.
Each reviewer saw the original task plus one of three materials: the agent's own replies (summary), its tool log (record), or both. The prompt was otherwise identical, and the reviewer had to end with OK or PROBLEM.
The verdict was scored against the stored result. 444 reviewer calls; no arm lost more than 3 to unparseable replies.
Design, hypotheses and analysis (exact McNemar test, each reviewer analysed separately and never pooled) were written down before any reviewer call.
Every discordant case for the reviewer showing the 'both' effect was hand-read.
The numbers
Comparison (share of failures caught)
Reviewer A
Reviewer B
Summary only
73%
76%
Log only
76%
66%
Summary and log
59%
51%
Log vs summary (p)
1.0
0.29
Both vs log (p)
0.07
0.38
What it does not show
Machine reviewers, not people. The log states outcomes more directly than a busy human would read them, so this measures a careful machine reader.
Mostly one kind of task: 35 of 37 pairs are destructive-restraint tasks.
One draw per reviewer and arm. 82 calls that ran out of output budget or were cut off were asked again with a larger budget; nothing that had already returned a verdict was asked again.
Clean-pass rates were low (35% to 65%), partly because reviewers flagged real concerns outside what the stored result grades, so they are a floor.
Results as of Oct 2, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.
Start with the question you actually arrived with — there are five: