/p/2026-10-02 · explainer
Paper explainer · 2609.38812 · Luo, Wei, Wang et al.

Its own review missed four wrong answers in ten.

Terminal agents check their own work almost without exception: once a complete candidate solution exists, 99.53% of runs go on to test it. The checking is where it falls apart. Across ten agents, 61.43% of the wrong candidates were caught and 49.36% of the caught ones were repaired, so roughly three in ten wrong solutions are recovered end to end and the rest are handed over with a clean self-report. Noticing and fixing are also separate skills that trade against each other by model, which is why a single self-correction score tells you almost nothing about where to spend.

01 · The measurement

Find the moment the answer was finished, then watch what happens next

gates passed so far

02 · The arithmetic

Two gates multiply, and the product is the number you actually get

A hundred wrong first answers, two gates

03 · The trade

The models that notice most, fix least

04 · The fix

Teach the checking, not the whole run

05 · Why I care

Work out how many wrong answers your self-check is actually stopping illustrative

Results

What the paper actually measured

What it does not show

In practice