Synthetic data is now a routine part of training and evaluation pipelines. One research result comes up in nearly every serious conversation about it: when models are trained on the output of other models, quality can degrade. This article explains what was shown, what it does not imply, and what it changes in practice.
In 2024, Shumailov and colleagues published "AI models collapse when trained on recursively generated data" in Nature. They trained models in generations, where each generation learned from data produced by the one before it, and tracked what happened to the learned distribution. They reported the effect in language models, variational autoencoders and Gaussian mixture models.
They describe two stages. In early model collapse, the model begins to lose information about the tails of the original distribution, meaning the rare and unusual cases. Overall performance can look stable at this stage, which is why it is easy to miss. In late model collapse, the model has converged on a narrow distribution that looks very different from the original data.
It does not say that synthetic data is harmful in general. The experiments concern a self-consuming loop in which real data is progressively replaced by model output. Training a model on a carefully designed, filtered dataset from a stronger model, while keeping real data in the mix, is a different situation.
Follow-up work points the same way. Gerstgrasser and colleagues (2024) compared replacing data at each generation with accumulating synthetic data alongside the original real data. In their settings, accumulation avoided the degradation that replacement produced. The practical reading is that the risk grows when real data disappears from the loop, not that synthetic data cannot be used.
Rare cases are often the expensive ones: an unusual refund request, an uncommon contract clause, a sentence that mixes Hindi and English mid-phrase, a user who words a question in a way no template anticipated. A generator that favours its most likely outputs will under-produce exactly these. Average quality metrics will not reveal it, because the common cases dominate the average.
Synthetic data works best as a supplement anchored to reality. If a vendor cannot tell you what real data their pipeline is grounded in, how rare cases are covered, and how the result performs on real examples, treat the dataset as unverified.
Synthetic datasets and evaluation benchmarks for teams shipping LLMs and AI agents, delivered with the evidence.