Model collapse is real — and it's why synthetic data quality is a buying criterion
A 2024 Nature paper proved AI models degrade when trained indiscriminately on their own outputs, and a newer finding that just 0.1% synthetic contamination is enough to hurt a model has kept the debate alive rather than settling it.
In July 2024, researchers from Oxford, Cambridge, and Imperial College published a paper in Nature that gave the AI industry a name for a problem it had been quietly worried about: model collapse. Their finding, stated plainly in the abstract: “indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear.” Feed a model enough of its own (or another model’s) output, uncritically, and it doesn’t just plateau — it forgets the rare, edge-case data that made it useful in the first place.
The debate HN has been having for two years
The paper’s Hacker News thread pulled 274 points and split cleanly into camps that still haven’t reconciled. Skeptics like commenter simonw argued the risk was overstated in practice: “I don’t think the ‘model collapse’ problem is particularly important these days. The people training models seem to have that well under control.” Others pushed back hard — commenter mvdtnz’s one-line reply, “And you base this on what? Vibes?”, sums up the counter-position: labs are secretive about training data, so nobody outside them can actually verify the risk is contained.
That argument hasn’t gone away — it’s sharpened. A more recent HN discussion surfaced a survey of 65+ papers on model collapse, citing a follow-up finding from Dohmatob et al. at ICLR 2025: per commenter Aedelon, “even 0.1% synthetic contamination in training data causes measurable degradation,” and, more troubling for anyone scraping the open web today, “no major dataset (FineWeb, RedPajama, C4) currently filters for AI-generated content.” Two years after the original paper, the tooling to actually discriminate synthetic from human-generated data at scale still doesn’t broadly exist.
Why this is a vendor problem, not just a research problem
This is exactly the gap that synthetic-data and data-labeling vendors are supposed to fill — provenance tracking, human-in-the-loop verification, and contamination-aware generation instead of naive scraping-and-retraining. It’s also why the category’s vendor landscape has been consolidating: NVIDIA acquired Gretel in March 2025 and folded its synthetic-data technology into the NeMo stack — Gretel no longer exists as an independently purchasable product, which is a real structural shift buyers evaluating this space need to know about, not a footnote.
The vendors still standing independently in this category — Mostly AI, Scale AI, Labelbox, Snorkel AI, and Surge AI — are increasingly differentiating on exactly the axis this research points to: how rigorously they can prove their synthetic or labeled data isn’t quietly degrading whatever gets trained on it, not just how much of it they can generate.
What it means for buyers
If you’re evaluating a synthetic-data or labeling vendor for anything downstream of model training or fine-tuning, “how much data can you generate” is the wrong first question — “how do you prevent contamination and verify provenance” is the one this two-year-old, still-unresolved debate says actually matters. We track these vendors and their current status on Crail’s synthetic data & data labeling category page, and if you’re comparing two of the larger players head-to-head, our Mostly AI vs Scale AI comparison is a reasonable place to start.