Thousands of Books Bought and Destroyed Daily to Train AI Sparks Outrage
The physical book industry is facing a threat amid the rapid development of artificial intelligence (AI). AI companies are now buying books in large quantities, not to read them, but allegedly to use them as material for training AI models.
This phenomenon has begun to be felt in a number of second-hand bookshops in the UK. Barter Books, a second-hand bookshop in Northumberland, usually sells around 2,000 to 3,000 books a week. However, the shop recently received an order for the same quantity in just one day from a company in Canada.
Several other bookshops in the region have also experienced a similar surge in demand. The shop owners cannot yet confirm what the thousands of books will be used for. But on the other hand, developments in legal cases involving AI companies provide strong clues about the possible purpose of the purchases.
AI companies have faced significant pressure in recent months from the publishing industry and authors over the use of copyrighted works to train AI models.
Anthropic, for example, recently reached one of the largest copyright infringement settlements in history. The company agreed to pay US$1.5 billion to more than 300,000 authors who filed a lawsuit against the AI company two years ago.
Meanwhile, OpenAI is still facing a number of copyright infringement lawsuits from various parties, including The New York Times, Encyclopaedia Britannica, comedian Sarah Silverman, and several non-fiction authors.
The plaintiffs accuse OpenAI of copying their books to train large language models (LLMs) and claim the company has enjoyed enormous financial benefits from exploiting copyrighted material.
Oddly enough, the Anthropic case has opened a loophole for AI companies to obtain training material legally. The court ruled that the company did not break the law when training AI using copyrighted works as long as the company purchased the books used. Another judge reached a similar decision in the Meta case last year.
This situation makes second-hand books a cheaper option for AI companies to obtain large quantities of source material. In addition, a number of publishers are likely reluctant to provide books directly to AI companies if they know the intended use.
The Anthropic case also revealed the existence of an internal programme called ‘Project Panama’. Anthropic described the project as an effort to destructively scan every book in the world.
In one memo that became evidence in the trial, the company explained that the training method used a code name because ‘we do not want it to be known that we are working on this.’
The method in question is destructive scanning. In this process, physical books are dismantled until their pages are separated. Each page is then fed into a high-speed scanner so that the information can be processed by the AI model.
After the scanning process is complete, the pages of the books are usually discarded, recycled, or destroyed into paper pulp.
Anthropic denied using rare books in the programme. In a statement to the BBC, the company said, ‘None of our data acquisition programmes purchase and destroy rare or antiquarian books.’
However, Anthropic is likely not the only AI company using this method. Other companies could apply destructive scanning under different rules.
Rare and antiquarian books could even potentially become highly valuable data sources for AI because they often contain information that is not available in ordinary books.
One recent book shipment, for example, was found to include at least one rare book. The shipment was later traced to an AI training facility owned by Amazon in Las Vegas that carries out destructive scanning of books for AI training purposes.
Most second-hand bookshops are also unable to identify who the buyers ordering books in bulk are. Orders usually come through marketplaces such as Biblio, which conceal the buyer’s identity.