The Unseen Cost of AI Training Data
The insatiable appetite of artificial intelligence models for training data has an unforeseen and, for some, a deeply troubling consequence: the destruction of physical books. While the digital realm offers vast repositories of text, the process of creating and refining large language models (LLMs) often involves sourcing data from physical copies. This can range from scanning existing digital archives to, in more extreme cases, acquiring and subsequently discarding physical books once their content has been digitized for training purposes. The implications are particularly stark for rare books, historical documents, and unique print collections that hold cultural and intellectual value far beyond their mere textual content. These are not just sources of data; they are artifacts of history, art, and human thought.
The primary driver behind this practice is the sheer scale of data required to train sophisticated AI models. Companies are constantly seeking to expand their datasets to improve model performance, reduce biases, and unlock new capabilities. When digital versions are scarce, incomplete, or not licensed for AI training, the allure of acquiring physical copies becomes a practical, albeit destructive, solution. This method, while efficient for data acquisition, bypasses the ethical considerations and long-term preservation efforts that have historically guided the curation of knowledge.

The Race Against Obsolescence and Destruction
The urgency behind the call to scan rare books stems from a dual threat. Firstly, the physical materials themselves are subject to degradation over time. Paper yellows, inks fade, and bindings crumble. Without proper archival care, these items are at risk of being lost to time regardless of AI development. Secondly, and more acutely, is the threat posed by AI companies’ data acquisition practices. If the primary incentive for possessing a rare book becomes its immediate digitization for training data, the physical object itself is often treated as disposable. This is a stark departure from traditional archival practices, which prioritize the preservation of the original artifact for future study and historical context.
This situation raises a critical question: what happens to the unique characteristics of a physical book that cannot be captured by a simple scan? Marginalia, annotations by previous owners, unique binding techniques, the texture of the paper, the smell of aged ink – these are all elements that contribute to the historical and cultural significance of a rare book. A digital scan, while preserving the text, strips away this tactile and historical dimension. The current approach by some AI companies risks erasing not just information, but the very context and materiality that make these artifacts invaluable.
A Call to Action: Digitization as Preservation
The response to this threat is gaining momentum in various corners of the digital preservation and open-access communities. The core idea is simple but vital: if AI companies are going to digitize these books for their own purposes, we must ensure that high-quality, archival-grade digital copies are created and made accessible to the public before the physical copies are destroyed or irrevocably altered. This involves leveraging existing scanning initiatives, supporting new ones, and advocating for open-access policies for digitized cultural heritage.
Libraries, archives, and independent digitization projects are at the forefront of this effort. However, the scale of the task is immense. Rare books are housed in institutions worldwide, and many are fragile, difficult to access, or lack the funding for comprehensive digitization. The current AI data-driven destruction model creates a perverse incentive: it highlights the value of these texts for data extraction, but simultaneously endangers their existence as physical objects. This paradox necessitates a proactive, global effort to digitize and preserve, not just for the sake of AI training, but for the enduring benefit of humanity.
The Technical and Ethical Hurdles
Digitizing rare books is not a trivial undertaking. It requires specialized equipment, skilled personnel, and significant financial investment. High-resolution scanners, careful handling protocols to prevent damage, and sophisticated metadata creation are all essential components of archival-quality digitization. Furthermore, copyright issues can complicate the public release of digitized materials, even if the physical books are old. However, these challenges should not deter the effort. Instead, they underscore the need for collaboration between AI companies, cultural institutions, and the public.
The ethical debate centers on the ownership and use of cultural heritage. Should AI companies be allowed to profit from data derived from artifacts they then destroy? Should there be a mandatory
