Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
An agent's commands are sampled from a distribution, so the one it runs is not always the best one it could have written — and a bad command, installing the wrong package say, changes the machine in ways that block every later step even though a better alternative was within reach. This paper spends test-time compute at the boundary between the model and the harness: sample several candidate commands from the same history, have a checker pick one, execute only that, and leave both the generator and the agent loop untouched. The sample count is not the lever, the checker is — with the agent's own 9-billion model checking, eight candidates moved first-try success from 50.0% to 54.8%, while a frontier checker on the identical candidates reached 68.0%, and checking at the level of single actions beat generating more whole runs at about a third of the token cost. If your agent executes commands, put a separate and stronger checker in front of each call rather than retrying whole trajectories.