News sites are blocking access to Internet Archive's Wayback Machine
Summary
This Marketplace Tech episode examines why some news publications are blocking the [[WaybackMachine|Wayback Machine]], the Internet Archive project that preserves snapshots of web pages. Andrew Deck of Nieman Lab says publishers fear that AI companies could use archived pages as a backdoor into copyrighted journalism, even though the source emphasizes that publishers did not point to direct evidence of this happening through the Wayback Machine. The episode’s core contribution is to connect AI Proxy Scraping Risk, Archive Access Tradeoff, and Internet History Fragility: protection against AI training-data extraction may also weaken the public record that journalists use for accountability reporting.
Key Claims
- The [[WaybackMachine|Wayback Machine]] sends crawlers to capture snapshots of web pages, making it a central example of Public Web Archiving.
- Some news publications are blocking those crawlers because they worry AI companies could access archived copies and train models on copyrighted news content.
- Andrew Deck says the publisher response is largely preemptive; the source reports no specific AI company or direct evidence that the Wayback Machine has already been used as the proxy-scraping route.
- Publishers’ fear is rooted in earlier large-language-model development, where journalism was used without upfront permission or compensation.
- The episode distinguishes the older concern that readers could bypass paywalls from the newer concern that generative AI companies could use archive snapshots for model training or commercial chatbot products.
- Deck describes the Internet Archive as facing a watershed moment because AI crawlers have weakened older honor-system norms around the open web.
- Public Web Archiving produces public value for readers, researchers, and journalists, but the episode argues that AI changes the perceived cost-benefit calculation for publishers.
- [[LibraryOfCongress|Library of Congress]] web archiving is presented as a possible state-supported model, but the episode immediately raises the risk that governments might try to alter or narrow the historical record.
- The source says the [[WaybackMachine|Wayback Machine]] has helped journalists track deleted or stealth-edited government pages, including during the [[DonaldTrump|Trump administration]].
- Mark Graham’s view, as summarized in the episode, is that publishers may hurt themselves when they turn away from libraries because journalists can lose evidence, documentation, and reporting tools.
- The closing promo connects the episode feed to How We Survive and Amy Scott, though the substantive source contribution is the archive, AI, and journalism-policy conflict.
Key Quotes
“without upfront compensation or permission” - Deck’s explanation of why publishers feel burned by earlier AI training practices.
“a watershed moment” - Deck’s description of the Internet Archive’s position under AI-crawler pressure.
“hurt themselves” - the episode’s summary of Graham’s warning about publishers turning away from libraries.
Connections
- Marketplace Tech and Stephanie Hughes - show and host context for the episode.
- Internet Archive, [[WaybackMachine|Wayback Machine]], Andrew Deck, Nieman Lab, Mark Graham, and [[LibraryOfCongress|Library of Congress]] - archive, reporting, and institutional cluster.
- AI Proxy Scraping Risk, Archive Access Tradeoff, Open Web Social Contract Erosion, Public Web Archiving, and Internet History Fragility - new concepts added by the episode.
- AI Content Licensing, Open Web Traffic Decline, and AI Answer Source Attribution - existing publisher-AI economics branch extended from answer traffic and licensing into archive blocking.
- Digital Preservation, Public Service Journalism, and AI Journalism Trust - preservation and reporting-trust branches extended by the Wayback Machine’s role in documenting changed or removed web pages.
- How We Survive and Amy Scott - Marketplace climate-solutions promo at the episode close.
Contradictions
- No direct contradiction found with existing wiki content.
- The source qualifies AI Content Licensing by showing that publisher concern can extend beyond direct AI-company crawling or negotiated deals into third-party archives that might preserve content outside current publisher access controls.
- The source qualifies Digital Preservation by adding an institutional conflict: preserving public internet history can become harder when rights holders treat open archives as potential AI training-data leakage.