/p/2026-09-18 · explainer
Paper explainer · 2609.18204 · Chen, Zhu, Zheng & Li

The trace moved the
threshold, not the eyesight.

Put one model in charge of checking another's work and show it the steps that were claimed. The worry is that a confident write-up makes the checker gullible. Over 4,551 judgments it did the opposite: catching real violations stayed at 99–100% no matter how elaborate the trace, while rejecting correct work climbed from 39% to 60%. The overseer did not get worse at seeing. It got harder to please.

01 · The setup

The same job, judged with four different amounts of story attached

Each task is a small recommendation with hard criteria: find the best French restaurant in Nashville that is under budget, good food, good service — all three must hold. Seven review excerpts sit underneath, one of which quietly kills a candidate (“$80 for a substandard meal”). Half the time the recommendation names the right restaurant; half the time it names the one the evidence disqualifies. The contradicting excerpt is always visible, so nothing is hidden from the overseer — this is a test of judgment, not of search.

The only thing that changes is how much procedural narrative comes attached: nothing, a four-step summary, a per-criterion walkthrough, or a walkthrough that explicitly claims to have verified the very criterion it violated. Step through them and watch the two rates move independently.

Interactive · two rates, four amounts of trace
0%50%100%

The blue bar never moves. Pooled across models, real violations were caught 99.3%, 99.5%, 100% and 100% of the time as the trace grew, and 98.1% of those catches named the criterion that was actually broken — the overseers were not fooled and were not guessing. The red bar is the whole paper.

02 · The decomposition

Accuracy hides two different things, and only one of them changed

Borrow a century-old tool from perception research. Any yes/no decision under uncertainty has two independent parts: sensitivity, how far apart sound and flawed work look to the decider, and the criterion, where along that scale it plants its flag and starts saying no. Accuracy is a blend of the two, which is why accuracy alone can drop for a reason that has nothing to do with the decider getting worse.

Drag the trace level. The two humps are what sound work and flawed work look like to the overseer; the gold line is the point past which it rejects. Watch which one moves.

Interactive · sensitivity and criterion, pulled apart
sensitivity (d′)
how far apart sound and flawed work look to this overseer — bigger is sharper eyesight
criterion (c)
where it plants its flag — more negative means it rejects on less
real violations caught
the hit rate — at ceiling in every condition
correct work rejected
the false-alarm rate — the number a longer trace moves

Sensitivity sits between 2.42 and 2.68 the whole way across — noise, not a trend. The criterion slides from −0.94 to −1.59. Fitted properly, each step up in trace detail multiplies the odds of rejecting correct work by 1.44 (95% confidence interval 1.28 to 1.63), and that survives every robustness check the authors ran: leaving any single task out, adding task-level fixed effects, and a thousand-replication bootstrap.

03 · The stated reason

Six in ten wrong rejections said the same thing: I cannot tell what this quote is about

The authors read the rejections. Two independent judges coded 934 stated reasons into categories and agreed almost perfectly (κ = 0.89), with two human annotators checking a stratified subsample. In the condition where the review excerpts showed a feature and a sentiment but not which candidate they described, about 60% of the wrong rejections said, in effect, that the overseer could not attach the evidence to the option under review — so it withheld approval.

The fix is a label, not a better model. Tag each excerpt with the option it concerns and that entire category disappears.

Interactive · what the excerpts are tagged with

Tagging cut the odds of a wrong rejection to 0.34 of what they were (95% confidence interval 0.17 to 0.67), worth 20 to 32 points of false-alarm rate at a fixed trace length — so this is not a verbosity effect in disguise. It is the cheapest intervention in the paper and the one most people can ship this week.

04 · The residual

Tagging removes the excuse; it does not remove the reflex

With every excerpt tagged, the attribution complaint drops to zero — and wrong rejections keep climbing with trace detail anyway. The stated reason just changes: the share of rejections coded as plain over-skepticism goes from about one in ten to about one in two. Something about a long, self-verifying narrative makes some overseers treat rigour as a reason to say no.

Not all of them. Toggle the trace and watch which models move.

Interactive · wrong rejections with every excerpt tagged
0%50%100%

Then there is the part that makes this hard to catch in production. The most trigger-happy overseer reported a mean confidence of 94.4 out of 100 when it was wrongly rejecting correct work, against 94.9 when it was correctly approving — a gap of half a point. It does not sound unsure when it is wrong. Two of the others drop 14 to 16 points on their bad calls, which is enough to route on.

05 · On your review loop

Whether the trace is worth showing is an exchange rate illustrative

Say you audit 1,000 items, half of which contain a real problem. Showing the overseer a detailed trace takes missed problems from 7 to 0 and takes wrongly rejected good work from 195 to 300. Whether that trade is good depends entirely on what a miss costs you relative to a needless rejection — a compliance breach and a re-run are not the same price.

Interactive · 1,000 audited items, your prices illustrative
no trace shown
detailed trace shown

The break-even lands near 15×. Below it — which is most product work, where a false rejection means a human re-reviews a fine answer — the trace is costing you. Above it, in the places where a miss is genuinely expensive, showing the trace still buys you very little, because detection was already at ceiling. The trace is doing almost no work on the axis you wanted it for.

Results

What the paper actually measured

What it does not show

In practice