Truth & Goodness
When Reputation, Power, and Public Criticism Collide
24 August 2026
The digital industry is returning to paper books in search of knowledge. It buys them, cuts them apart, and scans them to obtain texts written by human beings. In this sense, AI companies destroy books not to erase their content, but precisely to seize it. Among the copies bought in bulk, however, there may also be rare, out-of-print titles unavailable in digital archives. How many could disappear from circulation this way?
For years, AI models were trained mainly on content from the internet. Over time, however, the web has increasingly filled with automatically generated texts. Such material is often described as “AI slop”: mass-produced, uneven in quality, and often repeating information from other sources. Researchers and AI companies fear that models trained largely on machine-generated texts may, over time, produce worse and less varied answers.
Companies therefore began looking for “clean” data: texts written, edited, and published by human beings before generative artificial intelligence became widespread. As a result, books published before late 2022, when ChatGPT was released, are especially attractive.
Printed publications have significant value for AI labs. They contain texts created by humans and refined by authors and editors. Importantly, they hold specialised knowledge and varied forms of expression. These are harder to find in the chaotic resources of the internet.
The paradox is clear. The most digital industry in the world is returning to paper books because, online, it is increasingly difficult to separate texts written by humans from those generated by machines. Accordingly, the book, until now primarily an object of reading and cultural circulation, becomes, in this model, a strategic raw material for data.
The mechanism is simple. Companies, or intermediaries acting on their behalf, buy used, out-of-print, and niche titles. The books then go to facilities where their bindings are cut off and their pages are separated. Next, the sheets are fed through high-speed industrial scanners. The result is digital text that, once cleaned and processed, can feed internal data libraries.
The AI industry destroys books because so-called destructive scanning is faster and cheaper than methods that require pages to be turned by hand or scanned with specialised equipment. This method allows large batches of publications to be digitised efficiently. However, it ends with the destruction of the paper copy.
The best-documented case concerns Anthropic, the company behind the Claude chatbot. Declassified court documents show that the company bought millions of paper books, commissioned their cutting and scanning, and then stored the digital copies in an internal library used in work on AI models. Notably, one of the projects was called Project Panama. An internal memo stated:
Project Panama is our initiative to destructively scan all the books in the world. (…) We don’t want it known that we are working on this.
The discretion is largely a matter of public image. Intermediaries offering book-buying services to technology companies emphasized the importance of non-disclosure agreements and client anonymity. One company serving the book market put the calculation directly:
“An AI company destroys 2 million books” is not the kind of headline that wins public sympathy.
The anonymous nature of the purchases makes it difficult to determine the true scale of the practice. Court documents and journalistic investigations confirm destructive scanning in Anthropic’s case. They do not, however, allow us to state precisely how many books have been destroyed in this way across the industry. What is known is that intermediaries advertised the ability to fulfil orders ranging from thousands of copies to as many as 1 million.
After the issue drew public attention, some of them argued that they had not actually purchased books for AI training. Instead, they had merely been testing interest in such a service.
The strongest concerns come from antiquarian booksellers and used-book dealers. Bulk orders may include out-of-print, regional, specialist, or foreign-language publications. Not every such book is a unique artefact. Nevertheless, when purchases follow long ISBN lists, it is difficult to determine, case by case, how many copies of a given edition still exist and whether it is available in libraries. The problem is especially serious for titles that have not previously been digitised and have not entered public archives.
I don’t like that uncommon books are being pulped,
said a bookseller quoted by 404 Media, who specializes mainly in rare and small-print-run books. According to his account, in recent months his sales rose from around 20 books a week to hundreds.
Anthropic officially denied that it buys “rare or antiquarian” books, distinguishing them from “less common” titles. This did not reassure booksellers. Their fears concern precisely the possibility that, in a mass book-buying process, titles with very few surviving copies may disappear.
The AI industry is reaching for books because the internet, once imagined as an almost inexhaustible source of knowledge, is increasingly filled with machine-generated content. Meanwhile, language models need texts created by human beings.
AI companies destroy books not to wipe out their content, but to capture it and turn it into strings of characters that feed the next system. The question remains open: how many rare, out-of-print, or poorly accessible titles may vanish from circulation in this way?
Read thos article in Polish: Niszczą książki, by nakarmić AI. „Nie chcemy, by wiedziano”