/p/2026-10-01 · explainer
Paper explainer · 2609.35832 · Zhang

It fixed one answer and broke another.

“Now check your answer” is the cheapest reliability idea in the book, and it does two things at once: it repairs some wrong answers and it overturns some right ones. The accuracy number you read afterwards is their difference, which is why it can hide almost anything. Across 29 open-weight models on three tasks with no new evidence supplied, one 8-billion model gained 25.5 points on grade-school maths while turning 19.1% of its already-correct answers wrong — and on yes/no reading questions, a blanket second pass lowered accuracy in 20 of 28 model settings. Treat self-correction as a policy with a switch, not a free improvement.

01 · The problem

One accuracy number, two opposite effects

Intrinsic self-correction means asking the model to revise its own answer with nothing new added: no retrieval, no tool result, no human. Every item ends up in one of four boxes, and only two of them move the score.

Two rates describe the whole behaviour. Recovery is the share of initially wrong answers the second pass repairs. Harm is the share of initially correct answers it overturns. What you get is the difference, weighted by how much of each pile you had to start with — which means a model that is already good has far more to lose.

The identity behind every self-correction result

Notice what happens when you drag first-pass accuracy up. At 17% correct, a 35% recovery rate and a 19% harm rate is a large win. At 90% correct, the same two rates are a loss. The second pass has not changed at all; the population it is applied to has.

02 · The models

The headline gain and the damage are not the same story

Four real settings, all with the same protocol: generate an answer, hand the model the question and its own answer back with a revision instruction, take what comes out. Nothing external is added. Step through them and compare keeping the first answer against accepting every revision.

Keep the first answer, or accept every revision
model and task
0%50%100% correct

The first row is the one everybody quotes: a small model that gains twenty-five points from a second pass. It is real. So is the fact that nearly one in five of the answers it had right before, it gets wrong after. The second row is the same family at nine times the size on an easier task, and there the arithmetic flips: a hundred and twenty-two repairs against a hundred and fifty-nine breakages.

03 · How often it helps

On most settings, a blanket second pass is a downgrade

Across eighty-two model-and-task settings, revising everything lowers accuracy more often than it raises it on the reading task, and it is close to a coin flip on the other two. There is no task where the second pass is reliably free.

Settings where revising everything lowered accuracy
01428 settings

The obvious lever is the instruction. If the second pass is too eager to change its mind, tell it not to be. The controlled study ran five wordings on eleven matched models on the reading task — and every one of the five has a negative mean accuracy change. Wording moves the balance; it does not fix it.

Five revision instructions, same models, same items
how hard the prompt pushes the model to reconsider

The direction is consistent and it is worth internalising: the harder you push a model to find fault with its own answer, the more often it finds fault with a correct one. A prompt that says “change your answer if anything looks weak” is not a safety measure.

04 · The switch

Decide per request, using what the first pass already told you

If revising everything is wrong and revising nothing is wrong, the remaining option is to revise some things. The paper compares three runtime policies — keep the initial answer, always accept the revision, or invoke revision only when a signal from the first pass says it is worth it — and the gate is trained only on features you already have before the second call is made.

Gating is not a universal win either. It is chosen most often on the reading task and least often on the causal one, and where the first pass gives you no usable signal, a blanket policy beats it.

Which policy wins, per task
task
0

The token line matters as much as the accuracy line. On the reading task, the gate skips the second call on enough requests to cut revision output tokens by a median of 56.5%; on the causal task it skips almost nothing, and the saving is 2%. A gate that never fires is not a gate.

05 · Your own stack

What a second pass is worth on your traffic

Put your own numbers through the same identity. The starting point below is the small-model maths setting from section 02, and the two dials are the ones you actually control in production: how many requests you send for a second pass, and how good your trigger is at picking the wrong ones.

Gate coverage against gate precision illustrative
accuracy after the policy
second-pass calls you skip
0%50%100% correct

Two things fall out of playing with it. Precision is worth far more than coverage — a trigger that fires on a quarter of requests and is usually right beats one that fires on everything. And the moment your trigger’s precision drops below your first-pass error rate, you are doing worse than sending everything, which is a good reason to measure it rather than assume it.

Results

What the paper actually measured

What it does not show

In practice