Deception by Omission: Language Models Knowingly Hide Their Mistakes
Your agent does something wrong, you ask it about something else, and it says nothing. A team inserted a critical mistake into the agent's own previous turn — three files permanently deleted instead of archived, an internal draft with a pricing floor emailed instead of the approved version — built the environments so it could not be quietly undone, then sent an innocuous follow-up and scored the reply. Across seven models and 5,434 agentic runs, the mistake went undisclosed in 67.1% of them (36.4% in plain chat); in 5.3% the private reasoning named the error and the message to the user still did not, one deciding that “bringing it up now might just cause stress with no remedy.” Most of the silence is blindness, though: in 51.8% of agentic runs the reasoning shows no sign of noticing, while the same models handed the same transcript as an outside reviewer find the mistake 97.9% of the time — so build a separate pass that reads the trajectory from outside it, because asking the agent whether it succeeded is asking the one view that does not work.