scout.

a daily read of the ML and AI papers

FRI · 02 OCT 2026
3 papers

It could have run the better command

Three papers on who checks the agent's work: the better command it never ran, the mistakes its own review does not catch, and ten cheap judges standing in for an expensive one.

Today's pick
50.0% → 68.0%
of terminal tasks solved on the first try, once the harness samples eight candidate commands a step and a strong checker picks which one runs

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

An agent's commands are sampled from a distribution, so the one it runs is not always the best one it could have written — and a bad command, installing the wrong package say, changes the machine in ways that block every later step even though a better alternative was within reach. This paper spends test-time compute at the boundary between the model and the harness: sample several candidate commands from the same history, have a checker pick one, execute only that, and leave both the generator and the agent loop untouched. The sample count is not the lever, the checker is — with the agent's own 9-billion model checking, eight candidates moved first-try success from 50.0% to 54.8%, while a frontier checker on the identical candidates reached 68.0%, and checking at the level of single actions beat generating more whole runs at about a third of the token cost. If your agent executes commands, put a separate and stronger checker in front of each call rather than retrying whole trajectories.

61%, then 49%
of its own wrong solutions a terminal agent notices, and of those it notices, the share it manages to repair

Can Terminal Agents Trust Their Own Verification?

Telling an agent to check its work costs nothing and the agents already do it: across ten terminal agents, once a complete candidate solution existed, 99.5% of runs went on to test it. The checking is the part that fails — only 61.43% of the wrong candidates were caught, and of those caught only 49.36% were repaired, so around three in ten wrong answers are recovered end to end and the rest are handed over with a clean self-report. Noticing and fixing are also separate skills that trade off by model: one 35-billion open-weight model caught 71.53% of its own errors and fixed 25.35% of them, while a frontier model caught 49.60% and fixed 75.46%. Track detection and repair as two numbers rather than one self-correction score, and buy an outside check for whichever half is weak — in the paper's own fix, distilling only a stronger model's checking-and-recovery behaviour onto the student's own draft added 9.74 to 16.85 points of first-try success.

1/64 the price
of a frontier judge, for ten small open-weight judges stacked together, keeping a median 93% of its agreement with the gold labels

RAIM: Robust Aggregation of Inexpensive Models for Hallucination Detection

If you grade answers for faithfulness to their sources with a frontier model as the judge, the bill grows with your traffic, so this paper asks what a panel of ten small open-weight judges from unrelated model families gives up instead. Combined by a plain logistic regression — fitted out of fold so no judge is scored on its own data, and weighted so a judge that merely echoes a stronger one is discounted rather than counted twice — the panel keeps a median 93% of a frontier judge's agreement with the gold labels once chance agreement is discounted, gives up 2.9 points of balanced accuracy on average, and costs about a sixty-fourth as much per item. It is not a blanket win: across eight faithfulness sets the panel clearly beat the frontier judge on one and clearly lost on three. The transferable part is the cheap pre-check — are at least half your judges individually competent, and do their mistakes land on different items — which you read off a one-off 50-to-100-record labelled set before deciding to switch.