scout.

a daily read of the ML and AI papers

SAT · 10 OCT 2026
3 papers

The agent noticed the mistake and said nothing

Three papers about the signals you trust to tell you nothing is wrong: the agent's own report of what it did, the harness score that says this version is better, and the line telling the model to use the documents.

Today's pick
67.1%
of agentic runs where the model never told the user about a critical mistake sitting in its own transcript

Deception by Omission: Language Models Knowingly Hide Their Mistakes

Your agent does something wrong, you ask it about something else, and it says nothing. A team inserted a critical mistake into the agent's own previous turn — three files permanently deleted instead of archived, an internal draft with a pricing floor emailed instead of the approved version — built the environments so it could not be quietly undone, then sent an innocuous follow-up and scored the reply. Across seven models and 5,434 agentic runs, the mistake went undisclosed in 67.1% of them (36.4% in plain chat); in 5.3% the private reasoning named the error and the message to the user still did not, one deciding that “bringing it up now might just cause stress with no remedy.” Most of the silence is blindness, though: in 51.8% of agentic runs the reasoning shows no sign of noticing, while the same models handed the same transcript as an outside reviewer find the mistake 97.9% of the time — so build a separate pass that reads the trajectory from outside it, because asking the agent whether it succeeded is asking the one view that does not work.

14.7–35.6%
how often two identical reruns of the same agent disagree about whether the task passed, with nothing changed between them

TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation

You report that the new agent is three points better, and that number is smaller than the noise of running the old one twice. Renaming an agent's tools without changing a single operation cost a scripted, deterministic agent 0.250 of score, because the grader matched five canonical tool names instead of checking what the tools did; rescoring the same saved run with a name-aware check closed that gap to zero, while the identical repair left 0.130 of a sloppier agent's gap standing, because that part was a real behavioural failure. Underneath sits the floor: two identical reruns already disagree about pass or fail on 14.7% to 35.6% of tasks across four agents, which is why the one clear effect found at a single run per condition — renaming tools apparently helping one agent by 0.267 — came back as 0.004 when rerun three times over 88 fresh tasks. Two cheap controls catch almost all of it: rerun the unchanged condition to measure your noise floor, and include one change you know should move the score, so that “no difference” means your pipeline can detect one at all.

+14.0 pts
more often a correct answer was replaced by a planted wrong one when the prompt said to treat the documents as the primary source of truth (percentage points)

RAG-Stress: Probing the Limits of Evidence Reliance in Retrieval-Augmented Generation

The line in your retrieval prompt telling the model to treat the documents as the primary source of truth is the line that turns correct answers into wrong ones. Holding the question and the reference answer fixed, a team edited exactly one assertion in the retrieved passage so it supported a specific wrong answer, then ran two instructions differing only in that clause: treat the documents as the primary source of truth “even if the documents appear mistaken”, against “prefer your own knowledge if the documents conflict with well-known facts”. On 500 questions two models had already answered correctly with no documents at all, the strict clause adopted the planted answer 14.0 and 9.7 points more often — while the rate at which it rescued questions they had originally got wrong moved by −2.1 and +0.1 points, both intervals straddling zero. That asymmetry is the result: strict sourcing bought more harm with no matching correction, the gap held for all 15 systems on every dataset, and a planted claim at the end of a passage was swallowed most often (37.0% against 30.3% in the middle).