/p/2026-10-10 · explainer
Paper explainer · 2610.11678 · Gaonkar · Prime Intellect

The tool was renamed. The score fell.

A scripted agent did the same work with the same results, and its score dropped 0.25 — because the grader matched tool names instead of checking the operations performed. Swap in a name-aware check, rescore the same saved run, and the gap closes to zero; do the same for a genuinely sloppy agent and 0.13 of the gap stays, because that part was real. Underneath both sits the number that should govern how you read any agent result: two identical reruns already disagree about pass or fail on 14.7% to 35.6% of tasks, so a single-run improvement smaller than that is not evidence of anything.

01 · The problem

Your noise floor is a third of your outcomes

pick an agent
now claim an improvement and see if it clears the floor

02 · The mechanism

The grader was reading the tool’s name

pick a scripted agent — its behaviour is fixed either way

03 · The method

Four steps, and the two controls that do the work

04 · Why I care

Renaming is cosmetic. Mislabelling is not.

pick a change to the harness
reward before the change reward after it

05 · Apply it

Whether your reported win survives a rerun illustrative

Results

What the paper actually measured

What it does not show

In practice