An agent that can do the job is not an agent that can survive the job going wrong, and one number cannot tell you about both. This benchmark runs 36 enterprise workflows twice under one shared random seed — once clean, once with a single fault dropped at a named point in the call — so the two abilities get two numbers. Across 2,880 paired trials the models finished 83.5% of the clean runs but recovered only 46.7% of the runs they had already proved they could do, and at the cut point where the write commits and the confirmation is lost, retry-with-backoff ran the side effect a second time in 53.3% of trials. Having the agent read the far side before redispatching took recovery there from 43.1% to 75.0%; having the endpoint deduplicate took duplicates to zero.