Updated · 4 episodes · 3 shows · 4 source notes
AI Training Copyright Dispute
Definition
AI training copyright dispute is the legal and legitimacy conflict over whether copyrighted works can be used to train AI models without permission, payment, lawful acquisition, or a negotiated license.
Current Synthesis
The wiki now tracks the dispute across books, film, and music rather than treating it as one abstract “training data” issue. The strongest current synthesis is that legal risk depends on both what is learned and how the material was acquired: physical-book scanning, pirated-book downloads, film likeness and labor concerns, and music-label lawsuits all create different pressure points. The new All-In source sharpens the Anthropic branch by separating pirated LibGen acquisition from broader fair-use training claims and by adding output-training symmetry as a strategic problem for frontier labs.
Key Claims
- Acquisition path matters: buying, scanning, scraping, downloading pirated files, and licensing are not the same copyright posture.
- Fair-use arguments can coexist with settlements, lawsuits, and reputational pressure, so a single settlement does not settle the whole category.
- Creative-domain disputes connect training data to labor, likeness, market substitution, artist consent, and cultural legitimacy.
- Litigation and licensing can coexist across rightsholders, as shown most clearly in the music-source branch.
- AI labs’ objections to competitors training on model outputs can rebound against their own claims that training on human-created works should be allowed.
Evidence
- Book and AI-lab branch: Chip Stocks Crash, $20B Fund Margin Called, Frontier Labs: SLOW DOWN AI, Mamdani’s Grocery Stores adds the physical-book scanning and Google Books analogy through Anthropic, while The Fight Over Open Source AI, Anthropic’s $1.5B Payout, NYC Socialists: Evictions = Violence? adds the source-reported $1.5B Anthropic settlement, LibGen piracy distinction, and output-training hypocrisy argument.
- Film branch: EP277 对话贾樟柯(下):我没有背叛真实世界,我只是在寻找电影的新可能 uses Jia Zhangke to connect copyright with portrait rights, guild concern, and the legitimacy of AI cinema.
- Music branch: Can an AI music company make nice with human artists? uses Suno, Tatiana Cirasano, Universal Music Group, Sony Music, and Warner Music Group to show lawsuits and licensing deals evolving at the same time.
- Rights-holder strategy branch: Can an AI music company make nice with human artists? and The Fight Over Open Source AI, Anthropic’s $1.5B Payout, NYC Socialists: Evictions = Violence? both connect training-data disputes to collective bargaining, licensing, and whether direct competition changes the legal and commercial stakes.
Counterevidence & Qualifications
The source inventory does not provide final legal holdings for the whole AI-training category. The Anthropic discussion is source-scoped and distinguishes piracy from the separate question of fair-use training on lawfully acquired materials. The film source is primarily an ethical and creative-labor discussion, not a legal ruling. The music source shows that settlement by one label does not bind other labels, artists, publishers, or courts. Claims about model-output training, reviews, and knowledge diffusion remain analytical questions rather than settled law.
What Changed
- Migrated the page to
synthesis-v1with the original three sources preserved and the new All-In source appended. - Added the LibGen acquisition-path distinction to the Anthropic copyright branch.
- Added AI Output Training Symmetry as a separate strategic and legal-consistency problem.
- Compressed the prior connection pile into current synthesis, evidence groups, and related concepts.
Related Concepts
- AI Output Training Symmetry - consistency problem between training on human works and objecting to output distillation.
- AI Content Licensing - negotiated rights-clearing path that may reduce litigation pressure.
- Copyright Platform Conflict - historical platform-rightsholder conflict pattern.
- Digital Music Licensing - music-industry licensing analogy and active-rightsholder branch.
- User-Generated Content Copyright Risk - adjacent risk where the user supplies protected content rather than the model developer assembling training data.
- Creative Labor AI Backlash - labor and legitimacy response from creators whose work may be displaced or absorbed.
Sources
4 source notes across 3 shows
- Chip Stocks Crash, $20B Fund Margin Called, Frontier Labs: SLOW DOWN AI, Mamdani's Grocery Stores All-In with Chamath, Jason, Sacks & Friedberg
- EP277 对话贾樟柯(下):我没有背叛真实世界,我只是在寻找电影的新可能 Talk三联
- Can an AI music company make nice with human artists? Marketplace Tech
- The Fight Over Open Source AI, Anthropic's $1.5B Payout, NYC Socialists: Evictions = Violence? All-In with Chamath, Jason, Sacks & Friedberg