Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers
When one model audits another's work, the fear is that a long, confident write-up of the steps will talk the auditor into waving something broken through. Across 4,551 judgments from five reviewing models on 19 tasks where the contradicting evidence was always visible, that never happened — error detection held at 99-100% however elaborate the trace — but the threshold slid: each extra level of detail multiplied the odds of rejecting correct work by 1.44, taking wrong rejections from 39% to 60% while the reviewers' ability to tell right from wrong stayed flat. About 60% of those wrong rejections gave the same reason, that the reviewer could not tell which option a quoted excerpt was about; tagging every excerpt with the option it describes cut them by 20-32 points. Track your judge's false-alarm rate separately from its accuracy — the worst model here rejected good work 96% of the time and still reported 94 out of 100 confidence.