/p/2026-10-06 · explainer
Paper explainer · 2610.02503 · Paul and Nandy

Every component passed its own health check.

A retriever, a model and a few tools each work. The system they form does not, and the monitoring you have only sees the failures that crash. Across 150 production incident reports coded into 23 failure modes, 51% were silent — every check green, the answers wrong, a mean 4.2 days before anyone noticed, against 12 minutes for a crash. The same paper injects all 23 faults into a six-component pipeline and measures five guards you could wire in this week: circuit breakers that watch retrieval relevance instead of status codes, output quality gates, component isolation, semantic validators and runtime-typed boundaries. Three or more of them together took time to recover from 28.7 minutes to 8.4.

01 · The problem

Twenty-three failure modes, and the model is not one of them

where the failure starts

02 · Why nothing paged

Half of these failures keep every light green

the shape of the failure

03 · The mechanism

Switch the five guards on, one at a time

tap a guard to switch it on
a fault injected at the retriever, propagating right

04 · The trade

What each guard moves, and what it charges you

05 · In your own product

What the gate catches at your volume illustrative

Results

What the paper actually measured

What it does not show

In practice