UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents
The agent books the thing, the network loses the confirmation on the way back, the agent tries again, and now there are two bookings. A team built 36 workflows across payments, ticketing, git, storage, messaging, cloud, CRM and databases, then ran each one twice under one shared random seed — once clean, once with a single fault injected at a named point in the call — so that failing to recover is scored separately from never being able to do the task. Of the work two open-weight models could complete cleanly 83.5% of the time, only 46.7% survived a confirmation that never arrived, and retry-with-backoff repeated the side effect in 53.3% of trials; client-side idempotency keys bought 10.7 points of recovery and removed only 7.9 points of duplicates. Two things actually worked: having the agent ask the far side whether the write landed before redispatching, which took recovery at that cut point from 43.1% to 75.0%, and making the endpoint itself deduplicate, which closed 95.9% of the gap and took duplicates to zero.