Terminal agents check their own work almost without exception: once a complete candidate solution exists, 99.53% of runs go on to test it. The checking is where it falls apart. Across ten agents, 61.43% of the wrong candidates were caught and 49.36% of the caught ones were repaired, so roughly three in ten wrong solutions are recovered end to end and the rest are handed over with a clean self-report. Noticing and fixing are also separate skills that trade against each other by model, which is why a single self-correction score tells you almost nothing about where to spend.