Millions of books are reportedly being destroyed to train AI — here’s what Anthropic told Tom's Guide
The dismantling of classic literature for AI upgrades is disheartening
Today, the most popular AI models have become genuinely useful to many as a digital assistant. Yet, the companies behind them have drawn growing criticism. The rapid expansion of AI data centers has raised concerns about energy use, water consumption and the strain on local communities. At the same time, the huge demand for high-bandwidth memory (HBM) used to power advanced AI systems has contributed to supply shortages across the industry. Then there's the flood of "AI slop"—low-quality images, music, videos and text that increasingly fills social media feeds and search results.
Most people also know that AI models are trained on enormous collections of data. That information comes from across the internet, including websites, books, academic papers, images, audio recordings and videos.
Concerning books, recent reports from The Washington Post and 404 Media have exposed a disturbing trend — Anthropic (the company behind the Claude AI tool) has reportedly dismantled classic physical literature just so they could make the scanning process go faster in a bid to further train their AI models on their material. Here’s an explainer on why and how this trend has come to fruition.
Why and how this trend is happening
Anthropic’s name has been brought up in an operation referred to as “Project Panama.” The company reportedly used that codename to define a massive effort to acquire a vast amount of printed books and convert them into AI training data.
The AI company is said to have spent millions of dollars acquiring millions of physical books in bulk and proceeding to cut off the spines of said books so the individual pages could be more easily fed through high-speed industrial scanners. This process then led to those pages being converted into machine-readable text for machine learning—that information eventually became a part of a larger dataset used to train Claude AI. And due to those original books’ bindings being removed, they were disposed of as they were now considered nothing but trash.
The Washington Post’s report on this whole matter featured a quote from an internal planning document connected to Anthropic’s ultimate goal. “Project Panama is our effort to destructively scan all the books in the world,” that document stated. “We don’t want it to be known that we are working on this.” Clearly, word got out about the AI giant’s destruction of so much classic literature—this reveal resulted in a class action lawsuit brought against them by a collective of authors whose pirated work had been used to train Claude. This past July, Anthropic settled with those same authors at $1.5 billion.
Reuters reported that the court overseeing that case ruled that, while training AI on books is classified as fair use under copyright law, Anthropic violated the copyrights of several authors and their publishers by keeping 7 million pirated books of theirs in a central library. While companies such as OpenAI, Google and Meta have been embroiled in legal trouble regarding the use of books being used for AI training, Anthropic is the only one that’s been exposed thus far for actually dismantling physical books to train its models.
Get instant access to breaking news, the hottest reviews, great deals and helpful tips.
Another report from 404 Media alluded to a company called ISBNdb, a commercial book metadata and sourcing service that helps businesses locate books and bibliographic information, being used to find and purchase large quantities of physical books for sale to AI companies for their training efforts. Since that report went public, ISBNdb has scrubbed its webpages of any info related to training AI. Additionally, the company has denied ever buying, scanning or selling any books for AI training purposes and noted that its site was merely a "test of market interest."
Here’s what Anthropic had to say
A spokesperson for Anthropic tells Tom’s Guide, "Claude is trained on a mix of publicly available web data, commercially acquired datasets, and data we generate ourselves,” the Anthropic spokesperson told us. “Sourcing books is a widely used approach for training large language models across the AI industry. None of our data acquisition programs buy and destroy rare or antiquarian books."
Anthropic also provided more background on how they engage in the practice of using books to train AI:
- Sourcing books is a widely used approach for training LLMs across the AI industry.
- None of our data acquisition programs buy and destroy ‘rare’ or ‘antiquarian’ books. We buy books from regular commercial markets.
- Much of the chatter online seems to reference our recent Bartz settlement; as a reminder: We settled this case last summer, after the court found that training AI on books is allowed under copyright law [known as ‘fair use’] — a ruling that still stands.
- As Judge Alsup wrote: "Like any reader aspiring to be a writer, Anthropic's LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different."
- Bartz specifically concerns claims that Anthropic improperly downloaded materials from two online libraries — the LibGen and PiLiMi datasets.
- Anthropic never commercially released any model trained on the LibGen or PiLiMi datasets, and earned no revenue from any such model.
- We are pleased that more than 91% of authors and publishers covered by the settlement have claimed their share of the payment.
Bottom line
The visual of legendary books having their spines cut off, turned into nothing more than AI training data and later discarded is the stuff nightmares are made of. Rare and even out-of-print books may have been obtained by Anthropic for their past efforts, but we might never truly know until years later if the names of those demolished books become widely known. I’m sure no one wants to live in a world dominated by AI tools that have been trained on reading material that’s not readily available in their original physical form.
Follow Tom's Guide on Google News and add us as a preferred source to get our up-to-date news, analysis, and reviews in your feeds. Subscribe to Tom's Guide on YouTube and follow us on TikTok.
More from Tom’s Guide
- The smartest AI trick isn't a prompt — it's your Google Drive
- I found ChatGPT's secret menu — these 9 features change everything
- I tested ChatGPT Health for a week — these 7 prompts helped me get the best health insights

Elton Jones covers AI for Tom’s Guide, and tests all the latest models, from ChatGPT to Gemini to Claude to see which tools perform best — and how they can improve everyday productivity.
He is also an experienced tech writer who has covered video games, mobile devices, headsets, and now artificial intelligence for over a decade. Since 2011, his work has appeared in publications including The Christian Post, Complex, TechRadar, Heavy, and ONE37pm, with a focus on clear, practical analysis.
Today, Elton focuses on making AI more accessible by breaking down complex topics into useful, easy-to-understand insights for a wide range of readers.
You must confirm your public display name before commenting
Please logout and then login again, you will then be prompted to enter your display name.









