Logrine
Synthetic data Evaluation Blog About Trust Careers Book a scoping call
← All articles
Practice 11 October 2026 · 6 min read

What belongs in a dataset card for synthetic data

A dataset without documentation is a file of unknown origin. You cannot tell what it is for, how it was made, or when it will mislead you. For synthetic data the gap is wider, because the generation process is part of what the data is. A good dataset card closes that gap before the data is used.

Where the idea comes from

Gebru and colleagues proposed "datasheets for datasets" in work first circulated in 2018 and published in Communications of the ACM in 2021. Their framework is a set of questions covering a dataset's motivation, composition, collection process, preprocessing and labeling, recommended uses, distribution and maintenance. It was modelled on the datasheets that accompany electronic components, and it sits alongside Mitchell et al.'s "model cards" for documenting trained models. Dataset cards on public model hubs follow the same spirit.

The standard sections

What synthetic data adds

Generated data needs sections that real-world collections do not:

  1. Generation method. The generator models and versions, prompt templates, seeds and sampling settings, so the process can be understood and repeated.
  2. Seed provenance. Where seed examples came from, whether they contained personal data, and how they were de-identified.
  3. Synthetic-to-real ratio. How much of the data is generated and how much is real, since this affects the risks described in our article on model collapse.
  4. Filtering log. Each filter applied, its threshold, and how many records it removed.
  5. Quality evidence. The human audit sample size and results, judge calibration, duplication rate, diversity measures, and comparison against real data.
  6. Terms of use. The terms attached to the output of the generator models vary between providers, so record which apply to this dataset.
  7. Known limitations. Scenarios that are thin, languages that were reviewed less thoroughly, and biases that were measured but not removed.
  8. Version history. What changed in each release, so results from different versions stay comparable.

How to read one as a buyer

A short test: can you find the answers to these four questions in a minute? What real data is this anchored to? How was quality measured, on how many examples, and by whom? What is it known to be bad at? Could you regenerate or extend it? A card that cannot answer them describes a dataset you cannot yet trust.

The takeaway

Documentation is a quality control, not a formality. Writing down the limitations is also the clearest sign that the people who made the dataset measured it.

Sources
  • Gebru, T., Morgenstern, J., Vecchione, B., et al. "Datasheets for Datasets." Communications of the ACM, 2021 (first circulated 2018).
  • Mitchell, M., Wu, S., Zaldivar, A., et al. "Model Cards for Model Reporting." Proceedings of FAT*, 2019.
Keep reading
Model collapse: why synthetic data needs a real-data anchor
LLM-as-a-judge: three biases to control before you trust a score
Logrine

Synthetic datasets and evaluation benchmarks for teams shipping LLMs and AI agents, delivered with the evidence.

Product
Synthetic data Evaluation benchmarks Trust and data handling
Company
About Careers Contact hello@logrine.com
Resources
Blog Model collapse LLM-as-a-judge Dataset cards
© 2026 MaxxLabs. Logrine is a MaxxLabs product.
Bengaluru, India