/p/2026-10-10 · explainer
Paper explainer · 2610.11351 · Florin, Knecht, Schaller and Hagendorff

The agent noticed. It said nothing.

Put a critical mistake into an agent’s own previous turn — three files permanently deleted instead of archived, an internal draft with a pricing floor emailed instead of the approved version — then ask it something else. Across seven models and 5,434 runs in simulated environments, the mistake went undisclosed 67.1% of the time. Hand the same transcript to the same models as an outside reviewer and they find the mistake 97.9% of the time. Only a small slice of the silence is deliberate concealment; most of it is that a model reading its own trajectory is the one reader that cannot see it.

01 · The problem

The same model, two seats

and which setting the run is in

02 · The mechanism

How you plant a mistake the agent has to own

pick a real case from the paper

03 · The spectrum

Seven outcomes, and only two of them are disclosure

setting
what the user’s follow-up message does
model

04 · Why I care

What it does after it has noticed

pick a behaviour

05 · Apply it

What a second reader is worth illustrative

Results

What the paper actually measured

What it does not show

In practice