scout.

a daily read of the ML and AI papers

WED · 07 OCT 2026
3 papers

The retry charged the card twice

Three papers on what happens around the model rather than inside it: a retry that runs the payment a second time, a patch that passes every test and breaks the house rules, and a note typed under a paste that ends up welded into it.

Today's pick
53.3%
of plain-retry trials ran the external side effect a second time — a second charge, a second ticket, a second wire

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

The agent books the thing, the network loses the confirmation on the way back, the agent tries again, and now there are two bookings. A team built 36 workflows across payments, ticketing, git, storage, messaging, cloud, CRM and databases, then ran each one twice under one shared random seed — once clean, once with a single fault injected at a named point in the call — so that failing to recover is scored separately from never being able to do the task. Of the work two open-weight models could complete cleanly 83.5% of the time, only 46.7% survived a confirmation that never arrived, and retry-with-backoff repeated the side effect in 53.3% of trials; client-side idempotency keys bought 10.7 points of recovery and removed only 7.9 points of duplicates. Two things actually worked: having the agent ask the far side whether the write landed before redispatching, which took recovery at that cut point from 43.1% to 75.0%, and making the endpoint itself deduplicate, which closed 95.9% of the gap and took duplicates to zero.

43.1%
of the rules a project writes down were broken by patches that passed every test

Correct Code, Broken Contributions? Benchmarking Repository Policy Compliance for Coding Agents

A coding agent closes the issue, the tests go green, and the patch still breaks the house rules — the commit subject runs too long, the docstring convention is ignored, no test comes with the change. Researchers turned the contributor documentation of twelve well-known Python projects into 823 individually checkable rules, each with a small deterministic checker, then audited four models under two agent harnesses over 500 real issue-resolution runs, watching the whole session rather than only the final diff. Patches that were functionally correct still violated 43.1% of the rules that applied to them, and 50.3% of violations happened in intermediate steps that a review of the final diff never sees; the worst-kept category was the set of rules written specifically to govern AI contributions, at 18.0% compliance. Mounting every rule as a local file the agent can open lifts compliance by 8.75 points on average and gets the text into context in 99.6% of runs, and the best configuration still breaks about 30% — so the rules your merge actually depends on belong behind a deterministic check, not in a prompt.

81 of 100
code comments typed under a pasted snippet were pulled into the returned code by the strongest model tested

Can LLMs Separate Pasted Artifacts from User Speech? Absorption at Unmarked Prompt Seams

A user pastes a block of code or prose into your product and types a note underneath it — make this shorter, fix the off-by-one — and the model hands back the edit with the note welded into the text. A benchmark of 300 editing tasks, each run under six matched layouts of the same prompt across 20 models, measures how often: at a bare newline it runs from 7.7% to 66.7%, and inserting a blank line changes nothing in any model. Wrapping the pasted text in explicit start and end markers cuts it in 19 of 20 — the strongest model tested fell from 19.0% to 2.0%, and one added sentence saying that anything outside the markers is context rather than content took it to 0.0% — but the failure gets much worse when the trailing note looks like it belongs to what was pasted: a comment typed under code was absorbed more often in 17 of 20 models, and that same model absorbed none of 100 casual notes against 81 of 100 code comments. If your product splices user-pasted content into a prompt, delimit it and spend a sentence saying what the delimiter means.