AI Companies Destructively Scanning Books for Training Data
TECH

AI Companies Destructively Scanning Books for Training Data

38+
Signals

Strategic Overview

  • 01.
    Anthropic ran a secret internal initiative called Project Panama to buy millions of used physical books, cut off their spines with hydraulic cutting machines, scan every page, and discard the originals to build training data for Claude.
  • 02.
    Anthropic hired Tom Turvey, the former head of partnerships for Google Books, in February 2024 specifically to find a way to obtain 'all the books in the world' for the initiative.
  • 03.
    Internal planning documents directed employees not to publicize the effort, and details only became public through more than 4,000 pages of documents unsealed in the Bartz v. Anthropic copyright lawsuit.
  • 04.
    Vendors proposed converting between 500,000 and 2 million books over six months, and Anthropic spent tens of millions of dollars building the acquisition-and-scanning pipeline, contracted through Datamation Information Services.

Deep Analysis

Inside Project Panama: from bookshelf to shredder

Anthropic's secret internal initiative, code-named Project Panama, set out to acquire millions of used and rare books and convert them into training data for Claude. In February 2024 the company hired Tom Turvey, the former head of partnerships for Google Books, specifically to find a way to obtain 'all the books in the world.' [1]Internal planning documents were blunt about the goal: 'Project Panama is our effort to destructively scan all the books in the world.' [1]The process itself was industrial rather than careful - books were purchased in bulk, fed into a hydraulic-powered cutting machine that sliced off their spines, run through high-speed production scanners page by page, converted into searchable digital files, and the loose paper was then recycled. [1]More than 4,000 pages of documents detailing the operation only became public because they were unsealed in the Bartz v. Anthropic copyright lawsuit brought by book authors. [2]Anthropic's own planning materials show the company understood how this would look to the public: 'We don't want it to be known that we are working on this.' [3]Vendors ultimately proposed converting between 500,000 and 2 million physical books to digital form over a six-month window, and Anthropic spent tens of millions of dollars building out the acquisition-and-scanning pipeline. [1]

The court ruling that pays companies to destroy books

The court ruling that pays companies to destroy books
Scale of Project Panama and its legal aftermath, as disclosed in court filings and reporting.

The reason Project Panama could operate at all traces back to a June 2025 federal ruling. In Bartz v. Anthropic, Judge William Alsup found that training AI on legally purchased books was 'exceedingly transformative' fair use - but he also ruled that acquiring books through piracy, as Anthropic did via LibGen and the Pirate Library Mirror, was not fair use and amounted to willful infringement. [4]The distinction that matters for the destruction itself is subtler: Alsup's reasoning treated the physical book's destruction as evidence that the digital copy had 'replaced' rather than duplicated the original, which is part of what made the digitization legally 'transformative.' [5]Commentators have since argued that this logic quietly rewards companies for destroying their source material rather than retaining both formats, effectively turning book destruction into a legal safe harbor. [7]The ruling is described as having unleashed what one report called an 'unprecedented race' among AI companies to acquire and destroy books at industrial scale. [6]The financial stakes of getting the sourcing wrong are real - Anthropic ultimately agreed to a $1.5 billion settlement, the largest known copyright settlement in US history, to resolve claims tied to the pirated portion of its book acquisition. [8]

A hidden market of book brokers built for AI

Behind the individual scandal at Anthropic sits a less visible industry: middlemen who source books for AI labs in bulk while keeping the buyers anonymous. ISBNdb, identified by 404 Media, markets bulk sourcing of pre-2022 print books - between 1,000 and 1,000,000 per order - to AI companies, pitching them as guaranteed free of 'AI slop,' the machine-generated text now flooding the open web that risks causing model collapse if ingested during training. [9]ISBNdb wraps these deals in strict non-disclosure agreements so that AI buyers' identities, and the scale of what they're destroying, never surface publicly. Even the broker's own messaging acknowledges the reputational stakes: 'the optics problem is real,' since 'AI company destroys two million books' is not a headline that generates sympathy. [9]That anonymized, NDA-protected sourcing pipeline is a large part of why the scale of book destruction across the AI industry stayed hidden for so long, only surfacing through litigation discovery and investigative reporting rather than any company's own disclosure.

Not just Anthropic: an industry-wide pattern

Anthropic is the company with its internal memos exposed, but it is far from alone. Authors have separately accused OpenAI and Microsoft of breaching copyright in their own book acquisition for AI training; OpenAI has acknowledged downloading LibGen material but says it deleted the files before ChatGPT's release. [10]Google and Meta have also been named among the tech firms that went to significant lengths to obtain large troves of book data for training. [11]Google now faces a fresh 2026 lawsuit from major publishers, including Hachette, Cengage, and Elsevier, along with novelist Scott Turow, over its use of books to train Gemini - a case distinct from the 2015 precedent that found Google Books' search-snippet digitization to be fair use. [12]Meanwhile librarians, antiquarian booksellers, and cultural preservationists warn that destroying original books - especially rare or unique editions - erases historical artifacts permanently, even when the text itself survives in digital form. [13]Whatever the final tally of books destroyed across the industry, this is shaping up as a multi-company legal and cultural reckoning, not a single company's misstep.

Historical Context

2015
A federal appellate court ruled Google's book digitization for search and snippet display was transformative fair use, a precedent later distinguished from full-text AI training.
2021
Co-founder Ben Mann downloaded millions of books from LibGen, later found to be willful copyright infringement even though training itself was ruled fair use.
2024-02
Hired former Google Books partnerships head Tom Turvey and began Project Panama, its industrial book-buying and destructive-scanning operation.
2025-06-23
Judge Alsup ruled that training AI on legally acquired books is fair use, but that using pirated copies to build a 'central library' was not, sending that question to trial.
2025-09-05
Agreed to a $1.5 billion settlement with authors over pirated books used to train Claude, the largest known copyright settlement in US history.
2026-01-27
Reported on unsealed court documents detailing Anthropic's destructive scanning of millions of books under Project Panama.
2026-07-14
Major publishers including Hachette, Cengage, and Elsevier, and novelist Scott Turow sued Google over its use of books to train Gemini.
2026-07-21
Published an investigation identifying ISBNdb as a broker marketing bulk physical-book acquisition to AI developers seeking pre-2022, AI-slop-free training data.

Power Map

Key Players
Subject

AI Companies Destructively Scanning Books for Training Data

AN

Anthropic

Ran Project Panama, purchasing and destructively scanning millions of physical books to train Claude; separately downloaded millions of pirated books from LibGen and the Pirate Library Mirror, leading to a $1.5 billion settlement.

IS

ISBNdb

Middleman broker that sources pre-2022 print books in bulk (1,000 to 1,000,000 per order) for AI labs under strict NDAs, marketing the books as guaranteed free of AI-generated 'slop.'

JU

Judge William Alsup (U.S. District Court, N.D. Cal.)

Ruled that training AI on legally acquired books is fair use and that destructive scanning was justified because the digital copy replaced the print copy, while separately ruling that acquiring books via piracy was not fair use.

AU

Authors Guild / plaintiff authors

Brought the Bartz v. Anthropic lawsuit that forced disclosure of Project Panama's internal documents and won a $1.5 billion settlement covering over 480,000 works.

GO

Google and Meta

Named alongside Anthropic and OpenAI as companies that sought large troves of book data for AI training; Google now faces a separate 2026 lawsuit from major publishers over training Gemini on books.

LI

Librarians, antiquarian booksellers, and cultural preservationists

Warn that destroying original books, especially rare editions, permanently erases historical artifacts and call for greater transparency and non-destructive scanning alternatives.

Fact Check

13 cited
  1. [1] Anthropic's secret book-scanning operation revealed
  2. [2] Anthropic Project Panama scans and destroys books to train Claude
  3. [3] Anthropic Is Destroying Books
  4. [4] Federal judge rules in AI company's favor in landmark copyright case
  5. [5] What the evidence actually shows about AI companies destroying books
  6. [6] AI companies are reportedly shredding millions of books to train models
  7. [7] Why copyright law is quietly rewarding AI firms for shredding physical books
  8. [8] Anthropic settles with authors over pirated chatbot training material
  9. [9] AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop
  10. [10] AI companies' demand for printed books raises copyright concerns
  11. [11] AI firms destroying millions of rare books for training data, raising alarms over cultural heritage
  12. [12] Google faces another AI training lawsuit from major publishers
  13. [13] AI companies criticised for destroying rare books

Source Articles

Top 5

THE SIGNAL.

Analysts

Found that training AI on copyrighted books was exceedingly transformative fair use, and reasoned that destructive digitization was justified because the digital copy effectively replaced the print copy; separately found that acquiring books via piracy was not fair use.

Judge William Alsup
U.S. District Judge, Northern District of California

Internal characterization suggests Anthropic had legitimate paths to obtain books but initially chose piracy to avoid what Amodei called a drawn-out legal and business 'slog.'

Dario Amodei
CEO, Anthropic

Argues pre-2022 print books are uniquely valuable training data because they predate AI-generated contamination of the web, while acknowledging the reputational risk of the practice.

ISBNdb (company messaging)
Book-sourcing broker for AI companies
The Crowd

WTF AI companies are purchasing large quantities of used and rare books. Scanning their contents to train models. Then turning the originals to pulp. The sum of human thought, digitized and shredded. Every book that gets scanned disappears from the physical world permanently.

@@heyshrutimishra33402

Vean el gigantesco almacén de libros que la empresa de Inteligencia Artificial Anthropic obtuvo bajo el proyecto secreto Panamá, donde usó 2 millones de libros para entrenar a su IA y luego los destruyó para evitar denuncias de derecho intelectual. El capitalismo destruye y comodifica...

@@DaniMayakovski9358

Anthropic's Project Panama looks like this. the guys who decide if you're morally mature to use frontier models, who are sitting on BILLIONS of dollars, would rather do this than spend a little more and scan books like a normal human being.

@@Hesamation318

AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them at Incredible Scale, Even If Almost No Copies Remain

@u/Steap-Edit2800
Broadcast
AI companies illegally downloaded & destroyed millions of books

AI companies illegally downloaded & destroyed millions of books

La oscura verdad de la IA: millones de libros son destruidos para entrenarla | Maia Jastreblansky

La oscura verdad de la IA: millones de libros son destruidos para entrenarla | Maia Jastreblansky

Why Anthropic is destroying books #Vergecast

Why Anthropic is destroying books #Vergecast