Put one model in charge of checking another's work and show it the steps that were claimed. The worry is that a confident write-up makes the checker gullible. Over 4,551 judgments it did the opposite: catching real violations stayed at 99–100% no matter how elaborate the trace, while rejecting correct work climbed from 39% to 60%. The overseer did not get worse at seeing. It got harder to please.
Each task is a small recommendation with hard criteria: find the best French restaurant in Nashville that is under budget, good food, good service — all three must hold. Seven review excerpts sit underneath, one of which quietly kills a candidate (“$80 for a substandard meal”). Half the time the recommendation names the right restaurant; half the time it names the one the evidence disqualifies. The contradicting excerpt is always visible, so nothing is hidden from the overseer — this is a test of judgment, not of search.
The only thing that changes is how much procedural narrative comes attached: nothing, a four-step summary, a per-criterion walkthrough, or a walkthrough that explicitly claims to have verified the very criterion it violated. Step through them and watch the two rates move independently.
The blue bar never moves. Pooled across models, real violations were caught 99.3%, 99.5%, 100% and 100% of the time as the trace grew, and 98.1% of those catches named the criterion that was actually broken — the overseers were not fooled and were not guessing. The red bar is the whole paper.
Borrow a century-old tool from perception research. Any yes/no decision under uncertainty has two independent parts: sensitivity, how far apart sound and flawed work look to the decider, and the criterion, where along that scale it plants its flag and starts saying no. Accuracy is a blend of the two, which is why accuracy alone can drop for a reason that has nothing to do with the decider getting worse.
Drag the trace level. The two humps are what sound work and flawed work look like to the overseer; the gold line is the point past which it rejects. Watch which one moves.
Sensitivity sits between 2.42 and 2.68 the whole way across — noise, not a trend. The criterion slides from −0.94 to −1.59. Fitted properly, each step up in trace detail multiplies the odds of rejecting correct work by 1.44 (95% confidence interval 1.28 to 1.63), and that survives every robustness check the authors ran: leaving any single task out, adding task-level fixed effects, and a thousand-replication bootstrap.
The authors read the rejections. Two independent judges coded 934 stated reasons into categories and agreed almost perfectly (κ = 0.89), with two human annotators checking a stratified subsample. In the condition where the review excerpts showed a feature and a sentiment but not which candidate they described, about 60% of the wrong rejections said, in effect, that the overseer could not attach the evidence to the option under review — so it withheld approval.
The fix is a label, not a better model. Tag each excerpt with the option it concerns and that entire category disappears.
Tagging cut the odds of a wrong rejection to 0.34 of what they were (95% confidence interval 0.17 to 0.67), worth 20 to 32 points of false-alarm rate at a fixed trace length — so this is not a verbosity effect in disguise. It is the cheapest intervention in the paper and the one most people can ship this week.
With every excerpt tagged, the attribution complaint drops to zero — and wrong rejections keep climbing with trace detail anyway. The stated reason just changes: the share of rejections coded as plain over-skepticism goes from about one in ten to about one in two. Something about a long, self-verifying narrative makes some overseers treat rigour as a reason to say no.
Not all of them. Toggle the trace and watch which models move.
Then there is the part that makes this hard to catch in production. The most trigger-happy overseer reported a mean confidence of 94.4 out of 100 when it was wrongly rejecting correct work, against 94.9 when it was correctly approving — a gap of half a point. It does not sound unsure when it is wrong. Two of the others drop 14 to 16 points on their bad calls, which is enough to route on.
Say you audit 1,000 items, half of which contain a real problem. Showing the overseer a detailed trace takes missed problems from 7 to 0 and takes wrongly rejected good work from 195 to 300. Whether that trade is good depends entirely on what a miss costs you relative to a needless rejection — a compliance breach and a re-run are not the same price.
The break-even lands near 15×. Below it — which is most product work, where a false rejection means a human re-reviews a fine answer — the trace is costing you. Above it, in the places where a miss is genuinely expensive, showing the trace still buys you very little, because detection was already at ceiling. The trace is doing almost no work on the axis you wanted it for.