Amazon Is Slicing Spines Off Books to Feed AI Training

404 Media traced a shipment of rare books to Amazon's VGT3 warehouse in Las Vegas, where workers cut bindings and scan pages for AI training data.

An anonymous Amazon worker in Las Vegas described the workflow in plain terms to 404 Media: books arrive on pallets, workers feed them into an industrial binding cutter, the machine slices each spine off, and the loose pages are then run through fast scanners that flip them like cash-counting machines. The scanned sheets are dropped into six- to seven-foot cardboard bins called “shuttles” where they are mixed together and become unrecoverable. Duplicates and unwanted copies go in the same bins.

The worker is at Amazon’s VGT3 facility in Las Vegas, housed in the same building as LAS8, Amazon’s print-on-demand operation. VGT3 does not print books. 404 Media’s Emanuel Maiberg reported on 26 August that, in conversations with one of its employees, the warehouse’s purpose became clear: the books are being scanned for AI.

What the Worker Saw

The books arriving at VGT3 are not what most readers picture when they think of a corporate training corpus. The worker told 404 Media they had seen liquidated library books, pallets of Japanese-language books (some still factory-sealed), German and Russian volumes, academic materials from the University of London, and British parliamentary documents “presented to Parliament on behalf of Her Royal Majesty the Queen.” Some were brand new and still sealed; others looked rare.

The worker described the binding cutter in detail: “you slide the book into, remove your fingers from the area, and then you press a button, and it just comes down and slices it.” Between 20 and 25 scanners run at the same time, with attached screens where operators can watch each page capture. After scanning, the pages are loose sheets, “just a bunch of loose paper like just sheets thrown in there.” Unneeded books are “yeeted” into the same bins rather than repacked and returned to suppliers.

The worker also recounted the cover story they were first given: “At first what they were telling people is that it was for Kindles, but I just didn’t believe that.” Subsequent reporting confirmed the scans are destined for AI training.

How 404 Media Got There

404 Media’s earlier piece, published 17 August, worked backward from the books. 404 Media placed a tracking device in a shipment of rare books it suspected would be acquired by an AI company for training data, and followed the shipment around the country to its final destination. The tracker ended at the VGT3 facility in Las Vegas, whose team logo, according to 404 Media, is “a dinosaur, brandishing its teeth and with a book in its hands.”

The tracking investigation is what made the second, more detailed report possible. Once the destination was identified, 404 Media contacted VGT3 workers directly. Two outlets - TechCrunch and Futurism - confirmed the underlying findings the same day. Futurism added that, per the same VGT3 employee, some workers are “assigned to cut books” while others receive shipments and “scan the barcodes,” and that workers are apparently scanning the ISBNs of every book they process.

What Amazon Said

Amazon has not disputed the basic facts. The statement the company gave 404 Media, repeated in TechCrunch’s write-up, was: “Amazon purchases books through commercial channels to improve the products and services customers use.”

That response sidesteps the core of the story. It does not address which AI products the scans feed, what kinds of books are involved, why duplicates and unsold stock are being shredded rather than returned, or why library deaccessions and government materials appear to have ended up at the same station.

Why the Books Matter

Two things make book content especially valuable to frontier AI labs. First, large language models have largely exhausted the freely crawlable web. Anything that is not online - out-of-print titles, library-only research, foreign-language academic works - is now a scarce input. Second, books written before roughly 2022 are not contaminated by AI-generated text, which means training on them avoids the model-collapse problem that has begun degrading corpora scraped from the open web.

The collection pattern fits both pressures. TechCrunch noted the practice lines up with Anthropic’s $1.5 billion copyright settlement over allegedly pirated books. Futurism referenced a separate lawsuit in which a judge ruled that because Anthropic was turning the physical texts into digital ones and destroying the original copies, the practice was “transformative” and therefore did not violate copyright law. That ruling is the legal posture labs are increasingly relying on as the copyright landscape tightens around scraped text.

What This Means

The VGT3 story is a window into how the training-data economy now works in practice. The books are bought “through commercial channels,” as Amazon puts it. That includes bookseller listings, library deaccessions, foreign-language wholesalers, and estate sales - everything the worker described seeing on pallets. The downstream buyer is never named, because the warehouse handles scanning on behalf of an unnamed AI customer, not for Amazon’s own models specifically. The result is a supply chain that is hard to audit and even harder for the original rights holders to track.

There is a privacy beat here too, even though the books are physical objects. Library deaccessions often happen without itemized disclosure of where the discarded copies end up. Foreign-language and government documents arrive with no apparent provenance check. Workers at VGT3 say the process “changed every day,” which suggests the operation is still scaling up rather than running on a stable intake pipeline.

For readers who care about consent in training data, the practical takeaway is that the corpus is not just whatever is on the open internet. It is also whatever can be bought in lots, scanned fast, and shredded on a Las Vegas dock. A library copy sold off as discards does not lose its copyright status, and a UK Parliament document does not become public domain because it ended up on a pallet in Nevada.

The Bottom Line

A warehouse in Las Vegas is slicing the spines off books - including rare, foreign-language, academic, and government documents - and running the pages through scanners to feed AI training data. Amazon confirmed the practice. The downstream AI customer has not been named, and the original rights holders are largely not in the loop. The hard copy is destroyed in the process. That is the part of the AI training pipeline that does not show up in any model card.