“Now check your answer” is the cheapest reliability idea in the book, and it does two things at once: it repairs some wrong answers and it overturns some right ones. The accuracy number you read afterwards is their difference, which is why it can hide almost anything. Across 29 open-weight models on three tasks with no new evidence supplied, one 8-billion model gained 25.5 points on grade-school maths while turning 19.1% of its already-correct answers wrong — and on yes/no reading questions, a blanket second pass lowered accuracy in 20 of 28 model settings. Treat self-correction as a policy with a switch, not a free improvement.
Intrinsic self-correction means asking the model to revise its own answer with nothing new added: no retrieval, no tool result, no human. Every item ends up in one of four boxes, and only two of them move the score.
Two rates describe the whole behaviour. Recovery is the share of initially wrong answers the second pass repairs. Harm is the share of initially correct answers it overturns. What you get is the difference, weighted by how much of each pile you had to start with — which means a model that is already good has far more to lose.
Notice what happens when you drag first-pass accuracy up. At 17% correct, a 35% recovery rate and a 19% harm rate is a large win. At 90% correct, the same two rates are a loss. The second pass has not changed at all; the population it is applied to has.
Four real settings, all with the same protocol: generate an answer, hand the model the question and its own answer back with a revision instruction, take what comes out. Nothing external is added. Step through them and compare keeping the first answer against accepting every revision.
The first row is the one everybody quotes: a small model that gains twenty-five points from a second pass. It is real. So is the fact that nearly one in five of the answers it had right before, it gets wrong after. The second row is the same family at nine times the size on an easier task, and there the arithmetic flips: a hundred and twenty-two repairs against a hundred and fifty-nine breakages.
Across eighty-two model-and-task settings, revising everything lowers accuracy more often than it raises it on the reading task, and it is close to a coin flip on the other two. There is no task where the second pass is reliably free.
The obvious lever is the instruction. If the second pass is too eager to change its mind, tell it not to be. The controlled study ran five wordings on eleven matched models on the reading task — and every one of the five has a negative mean accuracy change. Wording moves the balance; it does not fix it.
The direction is consistent and it is worth internalising: the harder you push a model to find fault with its own answer, the more often it finds fault with a correct one. A prompt that says “change your answer if anything looks weak” is not a safety measure.
If revising everything is wrong and revising nothing is wrong, the remaining option is to revise some things. The paper compares three runtime policies — keep the initial answer, always accept the revision, or invoke revision only when a signal from the first pass says it is worth it — and the gate is trained only on features you already have before the second call is made.
Gating is not a universal win either. It is chosen most often on the reading task and least often on the causal one, and where the first pass gives you no usable signal, a blanket policy beats it.
The token line matters as much as the accuracy line. On the reading task, the gate skips the second call on enough requests to cut revision output tokens by a median of 56.5%; on the causal task it skips almost nothing, and the saving is 2%. A gate that never fires is not a gate.
Put your own numbers through the same identity. The starting point below is the small-model maths setting from section 02, and the two dials are the ones you actually control in production: how many requests you send for a second pass, and how good your trigger is at picking the wrong ones.
Two things fall out of playing with it. Precision is worth far more than coverage — a trigger that fires on a quarter of requests and is usually right beats one that fires on everything. And the moment your trigger’s precision drops below your first-pass error rate, you are doing worse than sending everything, which is a good reason to measure it rather than assume it.