Buying Old Books to Destroy Them: How AI Seeks Untainted Text
The best training material for artificial intelligence is sitting on a shelf. This phrase comes from ISBNdb, the largest book database in the world, which today sells a wholesale acquisition service for print volumes to IA labs to be scanned. There’s a sinister retribution: the sector that carries forward the promise of digitizing human knowledge is now paying to rake it in and shred it, one truckload of books at a time.
The crucial detail, which represents the value of the operation, is the print date: a book published before 2022 predates the era of large language models, so it cannot contain AI-generated text. The open web is now saturated with synthetic content produced by the same models, and training on that material exposes the so-called model collapse.
When a model is trained on the output of previous models, linguistic nuances flatten, systematic errors accumulate, and each generation turns out slightly worse than the one before. A print run from before 2022, on the other hand, is a fixed human archive that no one can secretly rewrite in retrospect.
The Race for Poisoning
Behind the demand for clean paper lies a second reason, and it’s the more interesting front. Authors have begun defending themselves with data poisoning techniques: inserting characters into the text that a person reads normally but that confuse a model, following tools like Nightshade designed for images. The idea is to make the work harmful for any model that ingests it, while remaining perfectly readable for humans.
ISBNdb cites in support an internal study from Anthropic, which states that it only takes between 250 to 500 expertly crafted documents, within a corpus of trillions of tokens, to implant a working backdoor in a model. It is not necessary to poison every book: a surprisingly small number suffices. Titles printed before 2022, written when these tools did not exist, circumvent the problem by definition.
Of course, ISBNdb has every interest in portraying web data as contaminated and its own volumes as the antidote. Model collapse and poisoning techniques are real phenomena, but their actual impact on large models remains quite nuanced.
The Image Problem
ISBNdb does not hide the fact that scanning these volumes almost always means destroying them. Workers cut the spine so that the loose pages flow into the scanner, faster and cheaper than the delicate alternative. This is why ISBNdb offers confidentiality to clients: strict NDA on every assignment, states the site, and the names of the clients are never disclosed.
This is obviously a refined issue of marketing and image: a company that destroys two million books is not something that generates sympathy. Better to speak of 'digital preservation of books'...