Compound AI System Reliability: A Failure Taxonomy and Resilience Pattern Catalog from 150 Production Incidents
When a system made of a retriever, a model and a few tools goes wrong, the fault is usually at the joins, not inside a part. Two researchers coded 150 production incidents — 97 from public post-mortems of twelve open-source projects, 53 from anonymised enterprise deployments — into 23 failure modes, and the dominant shape was silent degradation: in 51% of cases the system passed every health check while returning wrong answers, taking a mean 4.2 days to notice against 12 minutes for a crash. Injecting each fault into a six-component pipeline, 100 trials with and without each defence, they measured five patterns each worth about a day's work: circuit breakers tripping on retrieval relevance rather than HTTP status cut cascade depth from 3.8 components to 0.4, runtime-typed boundaries removed 92% of integration failures, and three or more patterns together took recovery from 28.7 minutes to 8.4. Add the output quality gate first — it catches 73% of silent degradation before a user sees it, for a median 120ms a request.