concept Updated 2026-08-07 Topics: Technology

Model Collapse

Model collapse is the failure mode in Kate Crawford: Mapping Empires where models trained repeatedly on synthetic outputs lose diversity, flatten minority patterns, erase outliers, and degrade toward lower-quality or noisier distributions. Kate Crawford also links the idea to model autophagy, where AI systems effectively consume their own outputs.

The concept turns synthetic data from a simple scaling solution into a risk that must be evaluated. It connects Frontier Model Scaling to Data Recipe Co-Creation: more data is not automatically better if the data is recursively generated, homogeneous, hallucinated, or detached from the human and physical variation the model needs to preserve.

174. 我们还能给算法当多久的品味老师?|对谈亚马逊AGI查晟 adds 查晟 / Cha Sheng’s training-practice version. He treats model collapse less as a ban on synthetic data than as a data-quality problem: generated examples can help if they carry information, pass filtering, and are evaluated, while low-information or unfiltered AI text can damage the training distribution.

Bytes: Week in Review - Micron’’s big earnings, Oracle’’s data center woes and “slop” is Merriam-Webster’’s word of the year adds a consumer-platform route into the same risk. The Marketplace Tech episode discusses whether large volumes of AI Slop could create feedback loops for future AI training, making platform quality and training-data quality part of the same problem.

Key Claims

  • Repeated training on generated outputs can narrow the distribution a model learns.
  • Minority patterns, rare cases, edge cases, and unusual styles are especially exposed when synthetic averages dominate.
  • Synthetic data may still be useful, but it needs grounding, filtering, evaluation, and diversity controls.
  • Synthetic data is not automatically harmful; its value depends on information content, filtering, and evaluation.
  • Media-scale AI Slop can become a data-quality problem if it enters future crawls.
  • Model collapse matters beyond aesthetics because high-stakes systems may depend on rare or outlier cases.
  • Public platforms can become part of the model-collapse risk surface if AI-generated material is published at scale and later recrawled as training data.

Connections