The world of publishing is facing a paradoxical reality: the physical book, which has served as a repository of knowledge for centuries, has become raw material for digital algorithms. Companies developing artificial intelligence have begun mass purchasing printed editions with a single goal — to train their models on them, after which the physical medium is destroyed.
As 404 Media found out, the procurement process often goes through intermediaries, allowing tech giants to hide their involvement in these deals. The volumes of trade are impressive: individual batches can range from 1,000 to 1 million copies. Buyers are not interested in genres, author names, or topics — for algorithms, the volume of text written by humans is what matters.
A New Market and Unexpected Players
Even services that previously worked exclusively with traditional book circulation are involved in this industry. For example, the ISBNdb database, which historically specialized in serving libraries and bookstores, now offers services for the mass purchase of literature for clients in the AI sector.
The used book market is reacting to this demand instantly. Sellers report a sharp spike in activity: if previously it was possible to sell about 20 books a week, now sales volumes have risen to several hundred copies. A similar trend is being recorded on major international platforms such as Alibris and Biblio.
The Technology of "Destructive" Digitization
The practice of using physical books to train neural networks has already received legal confirmation. In court materials regarding the company Anthropic, it was recorded that they purchased books, handed them to contractors for scanning, and then disassembled them into individual pages. This method allows for high-speed industrial scanning, which is significantly cheaper and faster than careful digitization that preserves the integrity of the publication.
The court ruled that the use of legally acquired books for training models is permissible under the principle of fair use. However, legal disputes in this field continue. Anthropic itself faced a lawsuit due to the storage of a library of 7 million pirated books. Furthermore, publishers have filed a separate lawsuit against Google, accusing the company of illegally using millions of copyrighted books to train the Gemini model.
The Paradox of Data Value
Experts note that printed books have become a valuable asset in the era of generative AI. Since they were written by humans, they do not contain "fluff" or artifacts characteristic of content created by other neural networks. This makes them ideal training material.
However, mass digitization raises alarming questions about the preservation of cultural heritage. After scanning, many copies are irretrievably destroyed. Particular concern is caused by the fate of rare and long-out-of-print editions: their physical copies disappear, while digital twins end up in closed databases accessible exclusively for the commercial training of algorithms.