scout.

a daily read of the ML and AI papers

SAT · 03 OCT 2026
3 papers

The harness decided which model won

Three papers on what your agent evaluation is actually measuring: the harness rather than the model, the same run twice, and a clean final answer with a broken trajectory underneath.

Today's pick
7.94 ahead → 30.16 behind
points separating the same two frontier models on the same 63 terminal tasks, once you swap the harness they run inside

Finding the Right Fit: Model-Harness Interactions across Agent Tasks

Choosing an agent stack means choosing two things — the model, and the harness that hands it tools, errors and context — and the second is usually treated as plumbing. Across 66 model-and-harness pairings the rankings do not survive the swap: on 63 terminal tasks one frontier model leads another by 7.94 points inside OpenHands and trails it by 30.16 inside the leaner PI, and for four of the five models the best harness changes from one benchmark to the next. What separates the harnesses is mundane — whether a failure comes back as something the model can act on: one has no shell timeout and left 34 runs silently blocked in a command that never returned, where another capped commands at 300 seconds and turned the same hang into a readable error. A model bake-off run inside someone else's harness is telling you about their harness, so run yours inside yours.

54%
of the spread in an agent's score that comes from re-running the identical configuration rather than changing anything about it

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

Agent evaluations now report cost and consistency alongside success, but they still hold the agent itself fixed and vary only the model. This team varied five parts of the configuration instead — how much task information the agent is given, reasoning, self-verification, time budget and the model — over 18,000 runs on four scientific coding tasks, and found about 54% of the spread in scores comes from running the same configuration again. Of the things you can actually change, what you tell the agent mattered most, ahead of model size, and it cut cost and improved calibration at the same time; more time only helped once the agent had enough information or a large enough model to use it. The sharpest result is about how you ask: pressing the agent in the prompt to verify its answer moved real reference-checking from 19% to 22% of runs, while handing it a verification tool roughly tripled that at every prompt level — build the behaviour into the system instead of requesting it.

162 of 175
trials that passed every final numerical check and still broke a rule somewhere in how they got there

Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills

If your agent eval asserts the final answer and nothing else, this is what it is not seeing. Running one business skill 240 times across two harnesses, three models and two versions of its written specification, the team scored the route as well as the result — which tools were called, with what arguments, in what order, and whether the database was left consistent. Of 175 trials that passed every applicable final numerical check, 162 (92.6%) carried at least one other detected deviation, and on a stricter seven-check definition of a correct final state 151 of 164 passing runs still violated a trajectory rule. The commonest roots were wrong arguments handed to the computation step, calling tools outside the allowed set, and skipping the read-back that confirms the write landed — none of which a final-answer assertion can see, and all of which are cheap to assert directly.