/p/2026-10-08 · explainer
Paper explainer · 2610.07753 · Lin, Cao, Luo, Chai, Lee and Hsu

The refund was right. It never read the order.

Agents judge consequential actions well on paper and take them badly in a loop, and this benchmark is built to keep the two apart: 656 cases over six operational domains, scored by a ledger that records which facts were established and when, so acting before you looked is caught even when the outcome happens to be right. On the same single-action cases, judging on paper runs 95–99% and executing with the evidence established first peaks at 52%. The failures sit upstream of the call — 21.7–62.9% of episodes stopped before the investigation finished and 37.0–66.9% fired before the required reads landed — while the action itself, once the evidence was in, succeeded 83.3–100% of the time.

01 · The problem

Judging the action and taking the action are different numbers

tap a model and harness

the six domains
02 · The mechanism

The evidence ledger: what was established, and when

tap an action

      
tap a read to perform it — or do not, and fire anyway

03 · The method

Five protocols, each adding exactly one thing

step through the protocols

04 · Where it breaks

Almost all of it goes wrong before the action runs

tap a measure to isolate it

05 · In your own product

What a missing read is standing in front of illustrative

Results

What the paper actually measured

What it does not show

In practice