/p/2026-10-02 · explainer
Paper explainer · 2609.39982 · Kang, Hachiuma, Zhang et al.

It could have run the better command.

A terminal agent’s next command is a sample, not a decision. Draw it again and you often get a better one — but nothing in the loop is looking, so the first draft is what reaches the shell, and a bad command changes the machine in ways no later step can undo. Mid-Harness puts the choice back: sample several candidate commands from the same history, have a checker pick one, execute only that, and leave the generator and the harness exactly as they are. The sample count is not the lever. Eight candidates judged by the agent’s own 9-billion model move first-try success from 50.0% to 54.8%; the same eight candidates judged by a frontier model reach 68.0%.

01 · The problem

The command that ran was a draw, not a decision illustrative

02 · The mechanism

Sample several, check them, run one

candidate commands sampled at each step

Plotted against the candidate count, the two groups separate immediately. The checker that is no stronger than the generator is flat; the checker that is stronger climbs with every candidate it is given.

First-try success against candidate count

03 · The comparison

How the checker is asked to compare costs more than what it concludes

how the checker is asked to compare

04 · The fix

You can teach the small model to judge without touching how it writes

05 · Why I care

Spend the budget on the step, not on the whole run illustrative

Results

What the paper actually measured

What it does not show

In practice