/p/2026-10-07 · explainer
Paper explainer · 2610.05622 · Sah, Sah, Jain and Sah

The retry charged the card twice.

An agent that can do the job is not an agent that can survive the job going wrong, and one number cannot tell you about both. This benchmark runs 36 enterprise workflows twice under one shared random seed — once clean, once with a single fault dropped at a named point in the call — so the two abilities get two numbers. Across 2,880 paired trials the models finished 83.5% of the clean runs but recovered only 46.7% of the runs they had already proved they could do, and at the cut point where the write commits and the confirmation is lost, retry-with-backoff ran the side effect a second time in 53.3% of trials. Having the agent read the far side before redispatching took recovery there from 43.1% to 75.0%; having the endpoint deduplicate took duplicates to zero.

01 · The problem

Finishing the job and surviving a fault are two different numbers

which number, and what it divides by

the eight domains
02 · The mechanism

Where the wire drops decides what recovery even means

tap where the fault lands
one mutating tool call, left to right

03 · The method

Four ways to retry, and only one of them asks

tap a recovery strategy

04 · The one that closed it

Move the deduplication to the far side of the wire

05 · In your own product

What a lost acknowledgement costs at your volume illustrative

recovery strategy

Results

What the paper actually measured

What it does not show

In practice