/p/2026-10-03 · explainer
Paper explainer · 2610.00917 · Li, Zhou, Teng et al.

The harness decided which model won.

An agent is a model and a harness — the program that hands it tools, decides what a failure looks like coming back, and cuts the run off when it stalls. Every comparison you have read varies the model and holds the harness still, as if it were a clean pane of glass. Across 66 configurations it is not: on the same 63 terminal tasks, Claude Opus 5 is 7.94 points ahead of GPT-6 Astra inside OpenHands and 30.16 points behind it inside PI. Neither model changed. The gap moved 38.09 points.

01 · The problem

A leaderboard is a ranking inside one harness

task set
read the ranking inside this harness

The full grid, for reference. Every cell is a mean reward over the same task set; the best harness in each row is marked.

All five models, all four harnesses
02 · The mechanism

A harness is what a failure looks like

03 · The spread

The same model, wrapped four ways

task set
Worst to best harness, per model

04 · The fallback that also fails

A model's own vendor harness is not reliably its best

Vendor harness score against the best outside harness
below the line — the vendor's own harness won above the line — an outside harness won

05 · Why I care

Your bake-off picks a model, then you change harness illustrative

task set your bake-off ran on
the harness you ran it inside

Results

What the paper actually measured

What it does not show

In practice