Agents judge consequential actions well on paper and take them badly in a loop, and this benchmark is built to keep the two apart: 656 cases over six operational domains, scored by a ledger that records which facts were established and when, so acting before you looked is caught even when the outcome happens to be right. On the same single-action cases, judging on paper runs 95–99% and executing with the evidence established first peaks at 52%. The failures sit upstream of the call — 21.7–62.9% of episodes stopped before the investigation finished and 37.0–66.9% fired before the required reads landed — while the action itself, once the evidence was in, succeeded 83.3–100% of the time.