/p/2026-10-03 · explainer
Paper explainer · 2610.01833 · Vo, Doshi, Sheinin

It passed. Then they looked at how.

An agent skill is graded the way everything else is graded: did the number come out right. Run one real business skill 240 times — two harnesses, three models, two ways of writing the same specification — and score the route as well as the result, and the two almost stop agreeing. Of 175 trials that passed every applicable final numerical check, 162 had already broken something on the way there: a tool outside the allowed set, arguments to the wrong scope, a write never read back. The answer was right in all 162.

01 · The problem

A passing grade, and a defect underneath it

what you count as passing

02 · The mechanism

Four things a route can get wrong while the answer stays right

03 · The findings

What was actually wrong, and how many causes it had

which breakdown

A single run trips several checks at once, because one wrong argument fails everything downstream of it. Walking the authored dependency graph collapses the cascade onto the failure that caused it.

Failed checks per run against root causes per run
failed checks the suite reports root causes left after attribution

04 · The other variable

The same skill, written down two ways

skill
harness

05 · Why I care

Which assertion to write first illustrative

the one assertion you add first

Results

What the paper actually measured

What it does not show

In practice