A retriever, a model and a few tools each work. The system they form does not, and the monitoring you have only sees the failures that crash. Across 150 production incident reports coded into 23 failure modes, 51% were silent — every check green, the answers wrong, a mean 4.2 days before anyone noticed, against 12 minutes for a crash. The same paper injects all 23 faults into a six-component pipeline and measures five guards you could wire in this week: circuit breakers that watch retrieval relevance instead of status codes, output quality gates, component isolation, semantic validators and runtime-typed boundaries. Three or more of them together took time to recover from 28.7 minutes to 8.4.