AI’s New Appetite for Books Is Putting Rare Titles at Risk

AI’s New Appetite for Books Is Putting Rare Titles at Risk AI’s New Appetite for Books Is Putting Rare Titles at Risk
AI’s New Appetite for Books Is Putting Rare Titles at Risk. Credit: New York Times.

Books are meant to be read, collected and preserved, but as Artificial Intelligence (AI) companies race to build more powerful models, some are reportedly being bought for an entirely different purpose: to be scanned, turned into training data and, in some cases, destroyed.

According to 404 Media, a growing network of intermediaries is helping AI companies source physical books on an industrial scale. Some reportedly handle orders ranging from 1,000 books to as many as one million, while keeping the identities of their clients confidential.

The demand comes from AI developers’ growing need for large volumes of high-quality, human-written material. Older books are particularly valuable because they were published before the explosion of generative AI and the spread of AI-generated material online, sometimes referred to as “AI slop”.

Advertisement

Read More

ISBNdb, a company known for its extensive book database, has reportedly moved into this market, helping AI companies source books in bulk and promoting pre-2022 print books as a cleaner source of training data. But sourcing the books is only the beginning.

When Books Become Training Material

Anthropic, the company behind Claude, has already been linked to the large-scale digitisation of physical books.

According to reports, court records from a 2025 case involving the company described how legally purchased books were stripped of their pages using a hydraulic-powered cutting machine and scanned with industrial imaging equipment. The process, known as “Project Panama,” was reported to have involved millions of books.

A United States federal judge, William Alsup, later ruled that scanning legally purchased physical books could qualify as transformative fair use, even when the originals were destroyed. The decision partly rested on the first-sale doctrine, which generally gives the lawful purchaser of a physical copy control over that particular copy.

The ruling did not give AI companies unrestricted rights to copyrighted works, however.

Anthropic also separately faced legal action over its alleged use of pirated digital books to train its AI systems, leading to an approved $1.5 billion settlement involving thousands of authors.

AI’s New Appetite for Books Is Putting Rare Titles at Risk
AI’s New Appetite for Books Is Putting Rare Titles at Risk. Credit: BBC.

Other major AI companies, including OpenAI and Meta, have also faced copyright cases over the use of books and other copyrighted material in AI training.

Now, what began as a way of digitising books is helping create a new market for them.

The Booksellers Feeling the Demand

One bookseller reportedly told 404 Media that his weekly sales jumped from about 20 books to several hundred after buyers he believed were linked to AI companies entered the market.

The orders were unusual enough to raise questions. His stock includes overseas and foreign-language books, as well as uncommon and out-of-print titles that can be difficult to replace.

The new business has been profitable, but the bookseller said he was uncomfortable with the possibility that some of those titles could be pulped after being scanned.

Similar suspicions have surfaced in the Netherlands, where 404 Media reported that rare booksellers have received bulk orders they believe may be connected to AI companies. Because intermediaries can purchase books for undisclosed clients, sellers often have no way of knowing who ultimately owns the orders.

That makes the scale of the practice difficult to determine and the fate of the books even harder to track.

For a common paperback, the loss of one physical copy may seem insignificant. For an out-of-print title with only a few surviving copies, the consequences are different.

The text can be extracted and preserved digitally. The physical book cannot.

AI’s New Appetite for Books Is Putting Rare Titles at Risk
AI’s New Appetite for Books Is Putting Rare Titles at Risk. Credit: New York Times.

What Could Be Lost?

The concern goes beyond copyright. A rare book carries more than its text; its annotations, illustrations, binding and provenance are part of its historical value. Once destroyed, those physical records cannot be recreated from a digital file or AI training dataset.

That is why some have called for non-destructive scanning, particularly for rare titles. Elon Musk, for instance, said he had instructed his AI team to preserve rare books in a library and scan them without cutting off their spines.

For AI developers, these books are valuable training material. For booksellers, they can be a lucrative new market. But for collectors, libraries and cultural institutions, the question is what happens when the race to preserve human knowledge in machines begins putting its physical record at risk.

The words may survive. The books will not.

Author

Share the Story

More From News Central TV

Advertisement

Keep Up to Date with the Most Important News

Weekly roundups. Sharp analysis. Zero noise.
The NewsCentral TV Newsletter delivers the headlines that matter—straight to your inbox, keeping you updated regularly.

×