Amazon destroys rare books to scan for AI training data
TECH

Amazon destroys rare books to scan for AI training data

27+
Signals

Strategic Overview

  • 01.
    A rare-book seller who listed about 1,000 books on the marketplace Biblio grew suspicious when a single buyer purchased the entire lot, and hid an Apple AirTag inside one volume so 404 Media could track where the shipment went.
  • 02.
    The AirTag led to a facility inside Amazon's LAS8 warehouse complex in Las Vegas, run under the internal name VGT3, where an employee confirmed the site's only function is scanning books, and workers reportedly cut off book spines to speed up the process, destroying each volume in turn.
  • 03.
    The scanned text is not converted into readable PDFs; it is fed into training data for Amazon's AI models, and Amazon confirmed it buys books "through commercial channels" but declined to confirm the AI training use.
  • 04.
    The practice echoes Anthropic's earlier book-scanning operation, which led to a $1.5 billion copyright settlement after a judge ruled that destructively scanning legally purchased books is fair use while acquiring pirated ones is not - giving AI labs a legal incentive to destroy books they buy outright.

Deep Analysis

Anatomy of the AirTag Investigation

A bookseller who had listed roughly 1,000 books for sale on the marketplace Biblio grew suspicious when a single buyer purchased the entire lot, and agreed to slip an Apple AirTag between the pages of one volume so journalists at 404 Media could follow the shipment [1]. The tracker led to a facility inside Amazon's LAS8 warehouse complex in Las Vegas, operating under the internal name VGT3 [2]. An employee who works at the site told reporters its sole function is scanning books rather than storing or reselling them: "I work at VGT3 here in Vegas, and all we do is scan books" [3]. Workers reportedly cut the spines off books to speed up the scanning process, physically destroying each volume as it is digitized [1]. Beyond scanning full text, workers also scan books' ISBN barcodes, and the resulting data is not converted into readable PDFs; it is used exclusively to help train Amazon's AI models [4]. Amazon confirmed to reporters that it purchases books through "commercial channels to help develop and improve the products and services our customers use," but declined to confirm or deny that the books are used for AI training specifically [5].

Why Old Books: The Model Collapse Problem

The push for physical books is driven by a scarcity problem: as the open internet fills up with AI-generated writing, models trained on that AI-generated text risk "model collapse," a degradation in output quality across successive generations, which is pushing labs toward pre-2022 human-written text instead [2]. Printed books are especially prized because much of their content was never digitized in the first place. As 404 Media co-founder Emanuel Maiberg put it, "Printed books are valuable as training data because a lot of the text they contain is not readily available on the internet" [6]. This is not a new theory inside the industry - internal documents from Anthropic's earlier book-scanning effort, known as Project Panama, reportedly show a co-founder theorizing that training on books could teach a model "how to write well" [6].

The Legal Loophole: Destruction as a Feature, Not a Bug

The specific act of destroying a book after scanning it traces back to a 2025 ruling in Bartz v. Anthropic, where Judge William Alsup found that destructively scanning a legally purchased print book to build a searchable digital library is fair use, because the resulting digital file simply replaces a print copy the company already owns - a use he called "transformative - spectacularly so" [7]. Critically, the same ruling found that Anthropic's separate acquisition of pirated books from shadow libraries was not protected in the same way, meaning the legality of destructive scanning hinges specifically on the book having been purchased legitimately first [7]. That distinction still carried a heavy price: Anthropic agreed to pay up to $1.5 billion to settle the underlying authors' class action over its pirated-book sourcing, described as the largest copyright settlement in U.S. history [8]. The Authors Guild is among the author advocacy groups whose copyright claims against AI companies over book-based training data resulted in that settlement, which is now a precedent shaping industry practice [9].

Rare Treasure or Dead Stock? A Contested Framing

The story's viral framing - "irreplaceable rare books destroyed for AI" - drew immediate pushback in online discussion. A large contingent of commenters argued that 404 Media never named the specific books in the shipment and cautioned against conflating "rare" with "valuable": much of what moves through bulk used-book channels is unremarkable dead stock that would likely have been pulped or shredded regardless of AI demand. A bookseller participating in the discussion echoed that point, noting that the volume of donated and secondhand books already exceeds what libraries and stores can absorb, so most of it gets shredded either way. A smaller, more vocal group dismissed the framing outright as "ragebait." That disagreement over how much is actually being lost does not resolve the underlying legal question, however: commenters on both sides of the debate pointed back to the Bartz v. Anthropic ruling as the specific mechanism that makes destroying the physical copy - valuable or not - a prerequisite for the digitization to qualify as fair use [7].

Historical Context

2025-06-23
Judge William Alsup ruled that Anthropic's destructive scanning of legally purchased print books for AI training was fair use, while acquisition of pirated books from shadow libraries was not.
2025-09
Anthropic agreed to pay up to $1.5 billion to settle the authors' copyright class action, described as America's largest copyright settlement.
2026-08-17
404 Media published its investigation tracking a rare-book shipment via AirTag to Amazon's VGT3 facility, revealing the destructive book-scanning-for-AI-training operation.

Power Map

Key Players
Subject

Amazon destroys rare books to scan for AI training data

40

404 Media

Independent journalism outlet that conducted the AirTag investigation and broke the story.

AM

Amazon

Operates the VGT3 scanning facility in Las Vegas; buys books in bulk, destructively scans them, and reportedly uses the text to train its AI models; confirmed purchasing but not the training use.

AN

Anthropic

Precedent case: ran a similar bulk book-destruction/scanning operation ("Project Panama"), later settled a landmark $1.5 billion copyright lawsuit (Bartz v. Anthropic) over pirated books used in training, while a court found destructive scanning of legally purchased books to be fair use.

RA

Rare/used booksellers (e.g., seller on Biblio)

Source of bulk book orders; one bookseller grew suspicious of a 1,000-book order and cooperated with 404 Media by hiding the AirTag; expressed concern that AI firms disregard books' historical and sentimental value.

AU

Authors / Authors Guild

Plaintiffs and advocacy groups pursuing copyright claims against AI companies over book-based training data, resulting in the Anthropic settlement precedent now shaping industry practice.

Fact Check

9 cited
  1. [1] We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility
  2. [2] AirTag reveals how Amazon destroys rare books for AI training
  3. [3] Amazon Caught Destroying Rare Books to Train AI
  4. [4] Tracking Rare Books Leads to an Amazon AI Training Facility
  5. [5] Amazon, once an online bookseller, is destroying rare books to train AI models
  6. [6] Now Amazon is destroying rare books to train its AI
  7. [7] Copyright and AI Collide: Three Key Decisions on AI Training and Copyrighted Content from 2025
  8. [8] The Bartz v. Anthropic Settlement: Understanding America's Largest Copyright Settlement
  9. [9] What Authors Need to Know About the Anthropic Settlement

Source Articles

Top 5

THE SIGNAL.

Analysts

Explains that printed books are valuable training data specifically because much of their text is not available online.

Emanuel Maiberg
Journalist/co-founder, 404 Media

Ruled that destructively scanning legally purchased print books to build a searchable digital library is fair use because the digital file merely replaces a print copy already owned, calling the training use "transformative - spectacularly so"; distinguished this from Anthropic's separate acquisition of pirated books, which was not found to be fair use.

Judge William Alsup
U.S. District Judge, Northern District of California (Bartz v. Anthropic)

Theorized that training AI models on books could improve models' writing ability.

Unnamed Anthropic co-founder (per Washington Post reporting on Project Panama documents)
Anthropic leadership
The Crowd

‼️ BREAKING: Journalists hid an Airtag in a rare book shipment and confirmed Amazon to be one of the buyers behind the bulk orders destroying rare-books to feed their AI. They tracked it to Amazon's LAS8 warehouse in Las Vegas, home to VGT3 (see logo below), a scanning operation

@@IntCyberDigest38886

JUST IN: Former online bookstore Amazon is buying old rare books and cutting them up to train AI.

@@remarks198

Hidden Airtag reveals Amazon is trashing rare books to train AI

@@arstechnica46

We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility / The shipment was found at an Amazon facility where Amazon scans and destroys books

@u/MarvelsGrantMan13621000
Broadcast
AI NEWS: An AirTag in a rare book led to an Amazon scanning site

AI NEWS: An AirTag in a rare book led to an Amazon scanning site

Daily AI News : Cutting Apart Rare Books to Train AI

Daily AI News : Cutting Apart Rare Books to Train AI

The most interesting thing in Product_Day 116_Rare books & AI

The most interesting thing in Product_Day 116_Rare books & AI

Amazon destroys rare books to scan for AI training data — AI News | Agentic Brew