Your agent asks to set invoice 4711 to paid. The application approves that, the call succeeds, and the row really does change — alongside a notification row a trigger inserted, or a second charge because the first response was lost and the call was retried. Approving the action and logging the aftermath both miss it, because nothing compares what actually persisted against what was authorised before the next step consumes that success. This runtime stages the writes inside a boundary it controls, collects the effects, and commits only on an exact match for this state, this plan revision and this one execution. On 206 public business tasks it kept 618 of 618 correct outcomes, rejected all 354 omission and 138 duplication variants an error-rejecting workflow accepted, and blocked invalid continuation in all 80 dependency topologies — for about 1.08 ms at the median.
Agent safety work almost always guards the action: is this tool call allowed, in this state, for this user. That check happens before execution, and execution is where the surprises live. A database trigger fires. A queue consumer picks the row up. The response is lost on the wire and the client retries. The approval was honest and the outcome is not the one that was approved.
Three shapes of this, from the paper. The left column is what the application authorised; the right is what was durable afterwards.
None of these is caught by an approval gate, and none is caught by an audit log either: the log records the extra row, after the workflow has already told its next step that the operation succeeded. The gap is between execution finished and success is propagated.
The runtime binds an approved outcome to three things at once — the current task state, the current plan revision, and exactly one execution occurrence, so an approval cannot be reused after a retry. It then runs the operation inside a boundary where every route that could persist anything is either covered by a declared profile or denied outright, collects the changes that would persist from audit records, and compares them against the approval at the profile’s granularity. Three outcomes: exact, diverged, or indeterminate when complete observation cannot be established. Only exact commits, and only exact releases the next step.
Switch a mechanism off to see which failure it was holding shut. The counts are the paper’s own ablations.
The ablation table is the useful part of the design. Each row is a mechanism you might be tempted to skip — and the specific way the system silently starts accepting wrong outcomes when you do. Result comparison alone accounts for 144 wrongly committed mismatches across the PostgreSQL and business-task sets.
The comparison baseline is the pattern most agent stacks already have: run the tool, discard the candidate if it raises, otherwise carry on. On 206 public business tasks covering shopping, customer support and travel, plus three database tasks, that workflow accepts every fault that does not raise — a write omitted, a write duplicated, a field changed that nobody asked about.
Step through the injected fault families. Every one is detected, and in every one the commit is withheld rather than logged.
Two results matter as much as the blocking. Clean executions are not disturbed — 618 of 618 correct outcomes still commit, so this is not a guard that buys safety by refusing work. And across 80 dependency topologies, every handoff matched durable state: when a step is withheld, nothing downstream is released on a success that did not happen.
The headline ratio looks alarming and is not: the median held-to-direct ratio is 4.35×, on direct task medians below one millisecond. In absolute terms the median added time is 1.08 ms. Comparison itself grows about linearly with the number of effects to check.
Certifying the execution boundary is the more visible cost, and it is still small: median 29.13 ms with certification against 13.29 ms without, a 2.19× multiple, of which the two certificates account for about 7 ms each.
You probably do not have a controlled execution boundary. You do have an agent that writes into systems with triggers, webhooks, queue consumers and retries — and a step that reads the tool’s success and carries on. Set your volume and how much of it touches something with side effects.
The cheap version needs no runtime. After a write, read back the rows and side effects the operation actually produced, diff them against the set the approval covered, and treat a mismatch as a failed step rather than an anomaly to log. Bind the approval to one execution so a retry cannot reuse it, and refuse to hand a success downstream until the diff is clean.