Buy, scan, destroy: AI firms are shredding millions of books to train their chatbots

AI companies are buying, cutting apart and scanning millions of physical books to train chatbots, according to a Washington Post report based on newly unsealed court filings. The documents reveal Anthropic's secret "Project Panama" and shed light ...

Agencies
Artificial intelligence companies are buying millions of physical books, cutting them apart, scanning every page and recycling the remains to train the chatbots powering today's AI boom, according to a report by The Washington Post, based on newly unsealed court filings in a copyright lawsuit against AI startup Anthropic. The filings also reveal how the race to build more capable AI models pushed companies to secure vast collections of books, setting off a wave of copyright lawsuits.

The court documents provide one of the clearest glimpses yet into the AI industry's search for high-quality training data. While Anthropic's internal book-scanning project is at the centre of the filings, separate lawsuits involving Meta, OpenAI and Google also highlight how leading AI companies sought access to millions of books to improve their models, although the methods differed across companies.

From bookshelf to chatbot

At the centre of the revelations is Anthropic's internal initiative called "Project Panama", which company documents described as an effort to "destructively scan all the books in the world." According to the report, the company spent tens of millions of dollars buying millions of used books, slicing off their spines, scanning every page and sending the remains for recycling to create training data for its Claude AI models. The planning documents also stated: "We don't want it to be known that we are working on this."


According to the report, Anthropic initially explored sourcing books from libraries and used bookstores, including New York's Strand Book Store, before eventually purchasing large batches from used-book retailers such as Better World Books and the UK's World of Books. A proposal from one of its scanning vendors said the company planned to digitise between 500,000 and 2 million books in six months, using industrial cutting machines to remove bindings before scanning the pages and sending the remains for recycling.

To lead the effort, Anthropic hired Tom Turvey, a former Google executive who helped create Google's Google Books project more than two decades ago, according to The Post.

The AI arms race

Books were viewed as a critical resource because they offered higher-quality writing than much of the internet. One Anthropic co-founder wrote internally that books could teach AI models "how to write well" instead of imitating "low quality internet speak," according to the report.
ADVERTISEMENT

Before launching Project Panama, Anthropic employees had also downloaded books from shadow libraries such as LibGen and Pirate Library Mirror, which host copyrighted material without permission. Anthropic has said it never trained a commercial AI model using the LibGen dataset and did not use Pirate Library Mirror to train any complete AI model.

The filings suggest Anthropic was not alone. According to The Post, Meta employees discussed using LibGen to obtain millions of books, with internal messages showing concerns over copyright risks. One engineer reportedly wrote, "Torrenting from a corporate laptop doesn't feel right," while another discussion referred to using rented servers to avoid the activity being traced back to the company. Meta has denied illegally distributing copyrighted works. OpenAI has acknowledged downloading LibGen but told a court it deleted the files before the release of ChatGPT. Google is also facing copyright litigation related to AI training data, with two major publishers recently seeking to join an existing lawsuit against the company.

The copyright plot twist

The legal battle over AI training data is far from settled. In June, a US judge ruled that using books to train AI models could qualify as fair use because the process is "transformative." However, the judge also found Anthropic could still face liability over how it acquired some of the books by downloading pirated copies.

Anthropic later agreed to pay $1.5 billion to settle claims related to the acquisition of books, while maintaining that the settlement concerned the method of acquisition rather than the legality of AI training itself. Most copyright cases involving AI companies, authors and publishers are still working their way through US courts, meaning the broader legal boundaries for training AI on copyrighted material remain unresolved.
Download
The Economic Times Business News App
for the Latest News in Business, Sensex, Stock Market Updates & More.
Download
The Economic Times News App
for Quarterly Results, Latest News in ITR, Business, Share Market, Live Sensex News & More.
READ MORE
ADVERTISEMENT

READ MORE:

LOGIN & CLAIM

50 TIMESPOINTS

More from our Partners

Loading next story
Business News › News › International › Global Trends › Buy, scan, destroy: AI firms are shredding millions of books to train their chatbots
Text Size:AAA
Success
This article has been saved

*

+