When Old Books Become AI Data

A New Opportunity for the Print Industry, or not?

Mysterious bulk purchases of secondhand books are raising an intriguing question: are AI companies turning printed books into training data? For printers, publishers and the wider graphic arts industry, the story highlights an unexpected new value for the printed page.

For years, the publishing industry has debated whether digital media would eventually make printed books less relevant. Now, artificial intelligence may be giving old books an entirely new kind of value.

Secondhand booksellers in the UK, Ireland, the US, Australia and elsewhere have recently reported a surge in unusual bulk orders. The purchases often involve hundreds or even thousands of books covering seemingly unrelated subjects — old technical manuals, biographies, foreign-language novels, outdated reference books and specialist publications that might otherwise have remained on shelves for years.

The identity and purpose of many of these buyers remain unclear. But one explanation is gaining attention: some of the books may be destined to become training data for artificial intelligence.

There is still no evidence proving that every mysterious bulk order is connected to AI. However, court documents involving AI developer Anthropic have already demonstrated that the idea of companies buying printed books, digitising them and using their content for artificial intelligence is very real. And that has important implications for publishers, printers, libraries and the broader print industry.

Millions of Books Turned into Data

Anthropic, the company behind the Claude family of AI models, became central to the discussion after documents from copyright litigation revealed details of an internal programme known as Project Panama. The company reportedly spent tens of millions of dollars acquiring millions of physical books.

Some were purchased in very large batches from established secondhand book suppliers. The objective was not to create a conventional corporate library. Instead, the books were converted into machine-readable digital text.

The process involved what is known as destructive scanning. Rather than placing each book carefully on a traditional book scanner, the binding can be removed mechanically and the pages separated. Loose pages can then pass rapidly through industrial scanning equipment, enabling huge volumes of material to be digitised.

Court documents showed plans capable of processing hundreds of thousands — potentially millions — of books within months. After scanning, the physical books could be sent for recycling. From an industrial perspective, it is a remarkably efficient conversion process: a printed product becomes a digital dataset. But it also changes the traditional relationship between printing and digital technology. The book is no longer merely competing with digital content. It can actually become the raw material from which new digital intelligence is created.

Why AI Wants Printed Books

Large language models require enormous quantities of written material to learn language, reasoning patterns, facts and writing styles. The internet has provided much of this data. But online content presents several problems. Web pages can be poorly written, duplicated, incomplete, inaccurate or extremely short. Increasingly, online material may also have been generated by AI itself. Books offer something different.

Professional books usually pass through authors, editors, proofreaders, designers and publishers before reaching the printing press. Their content is generally structured, coherent and significantly longer than typical web material. They also contain specialist information that may never have been freely available online.

An old engineering manual, chemistry handbook, historical reference work or technical printing guide might have almost no commercial value in the secondhand market. But if its information is absent from an AI company’s existing database, it could suddenly become useful. Perhaps most importantly, older printed books provide something that is becoming harder to guarantee on today’s internet: human-created content.

A book printed in 1985 clearly predates ChatGPT. A technical manual published in 1997 cannot have been generated by modern generative AI. That makes pre-AI printed material potentially valuable as a source of authentic human language.

The Problem of AI Training on AI

As generative AI becomes increasingly common, more internet content is being created or rewritten by artificial intelligence. This creates a challenge for future AI developers. What happens if new AI models are increasingly trained on material generated by older AI models rather than humans?

Researchers have investigated a phenomenon often described as model collapse, where repeatedly training models on synthetic data can gradually reduce the richness and diversity of the information they reproduce. The issue is more complicated than simply saying “AI cannot learn from AI.” Research suggests synthetic data can be useful when carefully managed and combined with high-quality real-world information.

But the debate strengthens the value of reliable human-generated material. And few sources provide a clearer historical archive of human writing than printed books produced before the generative AI era. In this sense, warehouses full of old books could become something resembling data reserves.

Why Are Buyers Choosing Strange Titles?

The pattern of reported purchases is particularly interesting. Booksellers have described buyers requesting combinations of titles with no obvious relationship: agriculture, engineering, biographies, historical publications, software manuals, fiction and books in different languages.

To a conventional bookseller, the selection can appear random. To an AI company building a dataset, however, it may make perfect sense. The objective may not be to acquire the most popular books. Popular titles are often already available in digital form. The real value could lie in books that are missing from existing databases.

An obscure mechanical engineering handbook from the 1970s could therefore be more valuable for data acquisition than a modern international bestseller. That reverses the conventional economics of publishing. A book that has almost no retail demand could still possess considerable informational value.

Printing’s Unexpected Role in the AI Economy

For the printing industry, this is perhaps the most fascinating part of the story. Print has traditionally been viewed as the final stage of content production. Authors create content, publishers prepare it and printers manufacture the physical product. AI may be creating a circular relationship. Printed material created decades ago can now be digitised and fed back into systems that generate new content.

That means the printing industry has inadvertently created an enormous offline archive of human knowledge — one that AI companies may now want to access. For book printers and publishers, the development also raises questions about future rights management.

If books possess value not only as products sold to readers but also as training material for AI systems, should publishers create new licensing models for machine learning?

Could archives of out-of-print publications become commercial data assets?

And might publishers eventually supply authorised, high-quality digital collections directly to AI developers rather than allowing physical copies to be bought and destructively scanned?

These possibilities could create an entirely new revenue stream for parts of the publishing industry.

But What Happens to the Physical Book?

There is also a less comfortable side to the story. Destructive scanning may be practical for mass-market books where thousands of copies survive. But the same approach becomes controversial when dealing with older, unusual or limited-edition material. A digital scan preserves the words and images. It does not necessarily preserve the physical object.

Paper stock, printing methods, binding, typography, handwritten notes, ownership marks and other physical characteristics can have historical importance. A book may therefore have two different values. For an AI developer, the valuable part may be its text. For a library, collector, printer or historian, the book itself can be an artefact. Once destroyed, that physical history cannot simply be recreated from a PDF.

Anthropic has said that its acquisition programmes do not target rare or antiquarian books for destruction. Nevertheless, concern among booksellers demonstrates the need for safeguards when physical publications become part of large-scale digitisation projects.

Copyright Questions Are Far from Settled

AI training has also triggered major copyright disputes. In the US, a federal judge ruled in 2025 that Anthropic’s use of lawfully acquired books for AI training could qualify as fair use under the circumstances of that particular case. However, the same litigation highlighted a critical distinction between legitimately purchased books and copies obtained from unauthorised online libraries.

Claims involving millions of pirated books eventually contributed to a settlement valued at $1.5 billion. The message for the industry is therefore not that AI companies have unlimited freedom to use books. Instead, the legal landscape is developing around questions of acquisition, copyright ownership, licensing, fair use and compensation. For publishers, authors and content owners, those questions will become increasingly important.

From Printed Page to Digital Intelligence

It is too early to conclude that AI companies are behind every mysterious order currently arriving at secondhand bookshops. Many buyers remain unidentified, while intermediaries linked to some transactions have denied knowingly sourcing books for destructive AI scanning. But the broader trend is undeniable. Printed books have acquired a new strategic relevance in the age of artificial intelligence.

For decades, the industry worried that digital technology would replace print. Instead, one of the world’s most advanced digital technologies may now be looking back at decades of printed material as a valuable source of knowledge. The dusty technical manual, forgotten novel or specialist reference book sitting on a secondhand shelf may no longer be simply an old publication waiting for a reader. It could also be data. And that puts the printed page in an unexpected position at the heart of the AI revolution.

Sources

BBC NewsSecondhand book sales are booming. Is it because of AI?
The GuardianReporting on unusual bulk book purchases in the UK and Ireland, August 2026.
The AtlanticInvestigation into mysterious international purchases of secondhand books and potential data-acquisition intermediaries, August 2026.
The Washington PostReporting based on court documents detailing Anthropic’s Project Panama and the large-scale scanning of physical books.
The Next WebReporting on Australian booksellers and concerns surrounding the acquisition and possible destruction of older books.
US Copyright OfficeReports and testimony on copyright, artificial intelligence training and emerging licensing markets.
Authors GuildReporting and documentation concerning the Anthropic copyright settlement.
NatureResearch on model collapse and the effects of recursively training AI models on model-generated data.

Exit mobile version