scout.

a daily read of the ML and AI papers

TUE · 06 OCT 2026
3 papers

Every component passed its own health check

Three papers on failures outside the model: at the boundaries between components, in a constraint stated forty turns ago, and in how the request happened to be phrased.

Today's pick
71% faster
recovery from an incident once three or more of the five resilience patterns are in place — 28.7 minutes down to 8.4

Compound AI System Reliability: A Failure Taxonomy and Resilience Pattern Catalog from 150 Production Incidents

When a system made of a retriever, a model and a few tools goes wrong, the fault is usually at the joins, not inside a part. Two researchers coded 150 production incidents — 97 from public post-mortems of twelve open-source projects, 53 from anonymised enterprise deployments — into 23 failure modes, and the dominant shape was silent degradation: in 51% of cases the system passed every health check while returning wrong answers, taking a mean 4.2 days to notice against 12 minutes for a crash. Injecting each fault into a six-component pipeline, 100 trials with and without each defence, they measured five patterns each worth about a day's work: circuit breakers tripping on retrieval relevance rather than HTTP status cut cascade depth from 3.8 components to 0.4, runtime-typed boundaries removed 92% of integration failures, and three or more patterns together took recovery from 28.7 minutes to 8.4. Add the output quality gate first — it catches 73% of silent degradation before a user sees it, for a median 120ms a request.

11.5%
of runs on the strongest model tested broke a rule that was still in force but had been stated many turns earlier

A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns

A rule your user sets early in a session quietly stops being enforced once it is far enough back in the history. The test: run the same task twice, once with the rule restated at the moment of action and once with it recoverable only from earlier turns, and count the cases safe the first way and unsafe the second. Over 412 matched instances across six tool domains, from calendar to filesystem to script execution, with histories of about 6,000 tokens over 56 to 160 turns, that gap was 11.5% on the strongest model tested and 27.8% on a 4B open model, and all seven models got worse as the history grew. The remedy is cheap enough to rebuild: extract each stated rule once into a small library with a scope, a trigger and the rule, re-render the matching ones into context whenever a task resumes, and put a deterministic check in front of the tool call that blocks a prohibited action and runs a missing prerequisite first. That check alone took the gap from 27.8% to 1.7%, beating every prompting and retrieval baseline they tried.

-22.4 pts
drop in an email agent's score when the same request is hinted at rather than asked directly (percentage points)

Lost in the Request: How Communication Variation Disrupts Retrieval and Action in Email Agents

Your eval asks each question one way; your users do not. Holding the needed information, the evidence and the correct outcome fixed and changing only the wording, an indirect phrasing — a hint or a stated need instead of a request — cost 22.4 points on one tool-using email benchmark, 10.7 on another and 3.2 on retrieval-augmented question answering, and it came out negative in six of six model-and-benchmark cells. Two mechanisms sit underneath: padding a request with chatter mainly breaks keyword retrieval, where recall of the right email fell from 95.7% to 44.6% between the least and most padded quarter of requests while a learned-sparse retriever held, whereas indirect and formal wording survive retrieval and surface as the agent leaving work undone — recall of the required tool calls fell 11.28 points with no significant rise in wrong ones, and formal phrasing took the share of tasks with no tool call at all from 0.4% to 12.2%. Borrow the method: generate style variants of your eval set, and score whether the work got done separately from whether the reply reads well.