Updated · 4 episodes · 3 shows · 4 source notes

concept Topics: Technology

AI Training Copyright Dispute

Definition

AI training copyright dispute is the legal and legitimacy conflict over whether copyrighted works can be used to train AI models without permission, payment, lawful acquisition, or a negotiated license.

Current Synthesis

The wiki now tracks the dispute across books, film, and music rather than treating it as one abstract “training data” issue. The strongest current synthesis is that legal risk depends on both what is learned and how the material was acquired: physical-book scanning, pirated-book downloads, film likeness and labor concerns, and music-label lawsuits all create different pressure points. The new All-In source sharpens the Anthropic branch by separating pirated LibGen acquisition from broader fair-use training claims and by adding output-training symmetry as a strategic problem for frontier labs.

Key Claims

  • Acquisition path matters: buying, scanning, scraping, downloading pirated files, and licensing are not the same copyright posture.
  • Fair-use arguments can coexist with settlements, lawsuits, and reputational pressure, so a single settlement does not settle the whole category.
  • Creative-domain disputes connect training data to labor, likeness, market substitution, artist consent, and cultural legitimacy.
  • Litigation and licensing can coexist across rightsholders, as shown most clearly in the music-source branch.
  • AI labs’ objections to competitors training on model outputs can rebound against their own claims that training on human-created works should be allowed.

Evidence

Counterevidence & Qualifications

The source inventory does not provide final legal holdings for the whole AI-training category. The Anthropic discussion is source-scoped and distinguishes piracy from the separate question of fair-use training on lawfully acquired materials. The film source is primarily an ethical and creative-labor discussion, not a legal ruling. The music source shows that settlement by one label does not bind other labels, artists, publishers, or courts. Claims about model-output training, reviews, and knowledge diffusion remain analytical questions rather than settled law.

What Changed

  • Migrated the page to synthesis-v1 with the original three sources preserved and the new All-In source appended.
  • Added the LibGen acquisition-path distinction to the Anthropic copyright branch.
  • Added AI Output Training Symmetry as a separate strategic and legal-consistency problem.
  • Compressed the prior connection pile into current synthesis, evidence groups, and related concepts.

Sources

4 source notes across 3 shows
  1. Chip Stocks Crash, $20B Fund Margin Called, Frontier Labs: SLOW DOWN AI, Mamdani's Grocery Stores All-In with Chamath, Jason, Sacks & Friedberg
  2. EP277 对话贾樟柯(下):我没有背叛真实世界,我只是在寻找电影的新可能 Talk三联
  3. Can an AI music company make nice with human artists? Marketplace Tech
  4. The Fight Over Open Source AI, Anthropic's $1.5B Payout, NYC Socialists: Evictions = Violence? All-In with Chamath, Jason, Sacks & Friedberg