The AI industry has a new appetite, and it’s an unsettling one: millions of physical books, bought in bulk, scanned, and then destroyed. To keep their names out of the headlines, major AI companies are reportedly routing those purchases through anonymous middlemen — a quiet supply chain built specifically to avoid the exact story you’re reading right now. For anyone tracking where the AI gold rush goes next, it’s a revealing look at what clean training data is really worth.
The Anonymous Book-Buying Pipeline
At the center of the practice is a class of intermediary services that let buyers place enormous orders while staying invisible. One such service, ISBNdb, reportedly facilitates orders of up to a million books at a time and keeps its clients anonymous, offering a “strict NDA on every engagement” and promising that buyers’ names are never disclosed. The result is a laundered supply chain: an AI lab gets the physical inventory it wants without its logo ever appearing on the invoice.
Booksellers, meanwhile, have described a sudden surge of orders for seemingly random rare and out-of-print titles — and a growing fear that they may be unwittingly profiting from the destruction of the very books they spent careers protecting. Understanding who you’re actually doing business with has never mattered more, a theme we’ve explored in our guide on choosing the right AI vendor.
Why Pristine, Pre-AI Text Is Suddenly Gold
The reason for all this cloak-and-dagger buying is almost poetic. As the open internet fills up with regurgitated AI “slop,” genuinely human-written text has become scarce and valuable. Physical books — especially those printed before large language models flooded the web — are among the last reliable troves of pristine, human-authored prose. To an AI company hungry for uncontaminated data, a warehouse of old books is a data mine.
The catch is that harvesting them at scale is destructive. Fast, cheap scanning typically means slicing the spine off a book so the loose pages can feed through a machine — a process that leaves the original physically ruined. Multiply that by millions of volumes and you get an industrial-scale shredding operation hidden behind the promise of smarter chatbots.
The Legal Gray Zone and the Cultural Cost
Legally, the companies have room to maneuver. A U.S. District Judge has ruled that digitizing legally purchased print books to train language models qualifies as fair use, reasoning that the process swaps a print copy for a digital one without creating extra copies or redistributing the text. That gives AI firms a defensible foundation — even as the ethics remain deeply contested.
The cultural cost is harder to quantify. Rare editions, once destroyed, don’t come back, and critics warn that irreplaceable pieces of print heritage are being fed into the grinder for marginal gains in model quality. It’s another reminder of how quickly AI is reshaping established industries, a shift we’ve tracked in how technology is changing business practices.
The Bottom Line
The scramble to buy and shred books lays bare an uncomfortable truth about modern AI: the technology’s hunger for quality data is now colliding with the physical world in ways that are wasteful and hard to undo. Whether regulators, publishers, or the companies themselves decide to draw a line, the episode is a preview of the resource battles the AI era still has coming.
