Logrine
Synthetic data Evaluation Blog About Trust Careers Book a scoping call
← All articles
Data quality 11 October 2026 · 6 min read

Model collapse: why synthetic data needs a real-data anchor

Synthetic data is now a routine part of training and evaluation pipelines. One research result comes up in nearly every serious conversation about it: when models are trained on the output of other models, quality can degrade. This article explains what was shown, what it does not imply, and what it changes in practice.

What the research found

In 2024, Shumailov and colleagues published "AI models collapse when trained on recursively generated data" in Nature. They trained models in generations, where each generation learned from data produced by the one before it, and tracked what happened to the learned distribution. They reported the effect in language models, variational autoencoders and Gaussian mixture models.

They describe two stages. In early model collapse, the model begins to lose information about the tails of the original distribution, meaning the rare and unusual cases. Overall performance can look stable at this stage, which is why it is easy to miss. In late model collapse, the model has converged on a narrow distribution that looks very different from the original data.

What it does not say

It does not say that synthetic data is harmful in general. The experiments concern a self-consuming loop in which real data is progressively replaced by model output. Training a model on a carefully designed, filtered dataset from a stronger model, while keeping real data in the mix, is a different situation.

Follow-up work points the same way. Gerstgrasser and colleagues (2024) compared replacing data at each generation with accumulating synthetic data alongside the original real data. In their settings, accumulation avoided the degradation that replacement produced. The practical reading is that the risk grows when real data disappears from the loop, not that synthetic data cannot be used.

Why the tails matter for business data

Rare cases are often the expensive ones: an unusual refund request, an uncommon contract clause, a sentence that mixes Hindi and English mid-phrase, a user who words a question in a way no template anticipated. A generator that favours its most likely outputs will under-produce exactly these. Average quality metrics will not reveal it, because the common cases dominate the average.

Controls we apply

  1. Anchor in real data. Seed generation from real or expert-written examples where they exist, and keep a real held-out test set that never touches the generation process.
  2. Use more than one generator. Mixing models reduces the chance that one model's habits define the dataset.
  3. Record provenance. Every record carries the generator, version and prompt template that produced it, so any slice can be traced or removed.
  4. Control and report the synthetic-to-real ratio. Treat it as a parameter, measured and written into the dataset card.
  5. Measure the tails. Track coverage of rare categories and scenarios explicitly, not just overall diversity.
  6. Evaluate on real data. A model trained on synthetic data should be judged on real examples before anyone trusts it.

The takeaway

Synthetic data works best as a supplement anchored to reality. If a vendor cannot tell you what real data their pipeline is grounded in, how rare cases are covered, and how the result performs on real examples, treat the dataset as unverified.

Sources
  • Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., Gal, Y. "AI models collapse when trained on recursively generated data." Nature, 2024.
  • Gerstgrasser, M., Schaeffer, R., et al. "Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data." arXiv, 2024.
Keep reading
LLM-as-a-judge: three biases to control before you trust a score
Deduplication and diversity matter more than dataset size
Logrine

Synthetic datasets and evaluation benchmarks for teams shipping LLMs and AI agents, delivered with the evidence.

Product
Synthetic data Evaluation benchmarks Trust and data handling
Company
About Careers Contact hello@logrine.com
Resources
Blog Model collapse LLM-as-a-judge Dataset cards
© 2026 MaxxLabs. Logrine is a MaxxLabs product.
Bengaluru, India