A terminal agent’s next command is a sample, not a decision. Draw it again and you often get a better one — but nothing in the loop is looking, so the first draft is what reaches the shell, and a bad command changes the machine in ways no later step can undo. Mid-Harness puts the choice back: sample several candidate commands from the same history, have a checker pick one, execute only that, and leave the generator and the harness exactly as they are. The sample count is not the lever. Eight candidates judged by the agent’s own 9-billion model move first-try success from 50.0% to 54.8%; the same eight candidates judged by a frontier model reach 68.0%.
Plotted against the candidate count, the two groups separate immediately. The checker that is no stronger than the generator is flat; the checker that is stronger climbs with every candidate it is given.