European booksellers are sounding the alarm after receiving massive, unusual bulk orders for thousands of obsolete non-fiction books that industry operators suspect are being purchased by artificial intelligence programmers for text training. According to Galway book retailer Kennys, the store recently secured an order for 5,000 second-hand volumes featuring eccentric titles ranging from decades-old driving manuals to guides on defunct financial schemes, mirroring similar bulk purchases flagged by vendors in Sweden and Germany.
The Anatomy of an AI Training Book Order
When Galway booksellers Tómas Kenny and Sarah Kenny received an online order for several thousand volumes through an undisclosed third-party buyer, they initially assumed it was a standard library acquisition. However, as staff began picking the stock from the warehouse upstairs, the reality of the purchase proved far stranger. According to Tomás Kenny, the selection was completely “scattergun,” focusing on materials that are “so obviously out of date and unusual” with no cohesive genre.
The eclectic mix included three-decade-old driving test manuals, guides to the Celtic Tiger-era Special Savings Incentive Account (SSIA), Internet Explorer for Dummies, and texts covering Caribbean history and medicine. Sarah Kenny noted that the bulk purchase pulled heavily from stock not normally open to the general public. While booksellers lack concrete forensic evidence linking the transactions directly to tech firms, the identical pattern reported across multiple European languages has heightened industry suspicions.
Did you know?
Booksellers in Sweden and Germany reported receiving the exact same style of bulk orders for obscure, language-specific second-hand texts prior to the Galway discovery.
How Second-Hand Books Feed Large Language Models
European booksellers suspect these eclectic shipments are consolidated at European collection points before being exported to the United States. Once stateside, operators reportedly slice the spines off the physical volumes and feed the pages through high-speed scanners to generate digital text datasets. These datasets are then utilized to train artificial intelligence models under broad legal interpretations of “fair use,” after which the physical books are destroyed.
This method circumvents traditional licensing agreements while securing massive quantities of human-written text. The reliance on older, out-of-print non-fiction provides AI models with diverse grammatical structures, historical reference points, and specialized vocabulary that cannot be easily scraped from modern web forums or sanitized news sites.
Copyright Precedents and the Growing Threat to Authors
The opaque acquisition of physical reading material intersects with an escalating global legal battle over intellectual property and AI training. In a historic precedent, a US federal judge granted final approval to a €1.3 billion copyright class-action settlement involving Anthropic. That lawsuit accused the artificial intelligence company of utilizing over 500,000 pirated books to train its Claude AI model, with eligible authors and publishers slated to receive roughly €2,600 per qualifying title.
Despite these legal victories, working authors face mounting professional displacement. Tomás Kenny pointed out that while physical bookstores have weathered previous technological disruptions—ranging from the invention of cinema and radio to the rise of the internet—the threat posed by generative AI targets the origin of the trade. Authors now face immediate concerns that generative tools will flood the market with algorithmic imitations, rendering human-written works commercially invisible and forcing publishers to cancel deals for fear of market saturation.
Frequently Asked Questions
Why are AI companies buying old books instead of using the internet?
According to industry observations, training advanced language models requires structured, long-form human prose, historical context, and specialized non-fiction that may not be readily available or legally accessible through standard web scraping.
What happens to the physical books after they are purchased?
Booksellers report that the acquired volumes are typically shipped to collection points and eventually to the United States, where the physical books have their spines sliced off for high-speed scanning before being destroyed.
Are authors being compensated for this use?
While some major copyright lawsuits have resulted in historic settlements—such as a €1.3 billion class-action involving Anthropic—many writers remain vulnerable to unauthorized AI mimicking and market displacement.
Join the Conversation: What are your thoughts on the impact of artificial intelligence on the publishing industry and physical bookselling? Share your perspective in the comments below, or subscribe to our newsletter for more investigative reports on the intersection of tech and culture.
Related reading