In recent weeks, headlines across major news outlets and social media feeds have warned of a grim new trend: artificial intelligence labs purchasing, scanning, and physically destroying thousands of rare books. The narrative presents a striking image of digital technology consuming and pulverising humanity’s printed heritage to feed mathematical models. It sounds like an alarming story, but it is also a fundamental misunderstanding of how data processing, copyright law, and archival preservation actually work. Rather than cultural erasure, what’s happening is a pragmatic attempt to bridge analogue records with digital systems.
Images of sliced book bindings and industrial paper shredders can look destructive without context. However, the process, known in document management as destructive scanning, is a standard mechanical solution to a logistical bottleneck. Overhead book scanners, which require manual page-turning, are slow and costly. Industrial sheet-fed scanners can process hundreds of pages per minute, but only after the binding of a book has been neatly severed. For high-volume data ingestion, slicing the spine off a paperback is simply the fastest way to achieve clean optical character recognition (OCR) without text distortion near the binding. This approach is far from unprecedented. Decades ago, long before modern language models existed, public institutions and private firms regularly implemented microfilming projects to compress room-sized paper archives into compact film rolls. Municipal offices and other administrative bodies routinely sliced, photographed, and recycled massive backlogs of paper records to keep archives functional. Rapid document ingestion today is the direct continuation of that same administrative practice. Much of the public anxiety stems from the claim that irreplaceable historical treasures are being intentionally destroyed. The operational reality, however, reveals a more complex and nuanced picture. Major tech firms are not targeting high-value, antiquarian collectables. The bulk shipments acquired by data processors include mass market overstock, obsolete technical manuals, expired encyclopedias, and uncirculated library surplus, material that municipal libraries routinely ‘weed‘ and discard by the millions each year to manage physical space. No report has established what share of these purchases is genuinely scarce material versus routine surplus, even among the most alarmed. Furthermore, national legal deposit laws protect the core intellectual record by requiring publishers to submit copies of every registered title to state registries, such as the Library of Congress in the United States. However, critics and booksellers raise a legitimate concern: automated bulk purchasing systems do not always distinguish between ordinary overstock and obscure, out-of-print titles. When automated brokers buy up large lots of niche non-fiction or regional histories, hard-to-find books with very few surviving physical copies can indeed end up in the industrial cutter. Conceding that obscure print material is occasionally caught in this net is essential, but it highlights a logistical casualty rather than a cultural crusade. Accidental collateral loss in a high-speed data supply chain is fundamentally different from the apocalyptic framing that AI companies are deliberately burning humanity’s collective heritage. It points to a failure of supply chain precision and cataloguing, not an intentional attempt to erase knowledge.
To understand why tech companies are acquiring physical paper rather than relying solely on web scraping, it is necessary to examine two primary factors: As generative AI tools flood the internet with automated text, the open web is becoming saturated with synthetic content. Machine learning researchers face a technical hurdle known as ‘model collapse‘, a phenomenon where AI models trained on AI-generated text gradually degrade in quality, logic, and coherence. To build reliable systems, developers require massive volumes of human-written, structured language. Physical books published prior to 2022 represent one of the clean, uncorrupted reservoirs of human thought, free from synthetic noise. Digital copyright law remains a complex legal battlefield. Downloading pirated digital libraries off the internet exposes tech companies to massive legal liabilities and copyright lawsuits. However, under long-standing property rules like the Doctrine of Exhaustion (known in the United States as the First-Sale Doctrine), purchasing a physical copy of a work grants the buyer explicit ownership rights over that specific physical item. By purchasing physical print on the open market, digitising the text, and discarding the paper original, companies argue they are exercising their right to transform physical property they legally own, creating a more legally defensible operational trail than web scraping. Beyond headlines and court rulings, the trend reflects how society’s view of written knowledge is shifting. Historically, libraries and archives treated books as public infrastructure and cultural assets to preserve, curate, and pass down across generations. In the industrial AI pipeline, books are evaluated primarily as raw material: units of high-quality training tokens to be bought by weight, extracted, converted into mathematical vectors, and discarded. This does not mean human heritage is being erased, but it does expose the inefficiencies of current digital policy. When tech companies must buy and scan physical paper just to get clean training data, it signals that our legal and archival structures have failed to create transparent, licensed pathways for sharing digital knowledge. International governance bodies and public policymakers have an opportunity here. Rather than restricting physical scanning, the focus could be on building open, ethical, and standardised digital repositories that make physical workarounds unnecessary. Throughout history, from the destruction of ancient libraries to wartime bombardments of national archives, true book destruction was designed to erase knowledge and extinguish cultural memory. Industrial digitisation operates under the opposite premise. It deconstructs the physical paper vessel to capture, index, and preserve the underlying data digitally. Ultimately, the surge of alarming headlines says less about the reality of technology operations and far more about the mechanics of modern digital media. Dramatic narratives depicting artificial intelligence on a relentless quest to consume human culture serve as reliable magnets for online engagement and outrage. This recurring, overused trope relies on sensationalism to generate clicks, obscuring practical realities behind emotional imagery. Instead of viewing this process through a lens of apocalyptic fear, tech policy analysis benefits from a calm, grounded approach. When we strip away the sensationalist framing, we can see this trend for what it really is: an awkward, transition-phase workaround to digitise human knowledge, reminding us that understanding complex technology requires looking beyond the headline. Author: Slobodan KovrlijaThe mechanical reality behind the headlines
Unpacking the ‘rare book destruction’ coverage
The technical & legal drivers behind the trend
1. The search for human-authored text
2. The legal framework of physical property
Knowledge as infrastructure vs. commodity
Preserving content over media