AI companies destroy physical books – let's scan rare books before it's too late
Cider9986
222 points
151 comments
August 21, 2026
Related Discussions
Found 5 related stories in 52.2ms across 4,128 title embeddings via pgvector HNSW
- AI companies destroy physical books – let's scan rare books before it's too late darccio · 703 pts · August 21, 2026 · 99% similar
- AI companies are shredding rare books anon373839 · 757 pts · July 27, 2026 · 74% similar
- Amazon, which started off selling books, is destroying rare texts to train AI rzk · 91 pts · August 17, 2026 · 71% similar
- Apple AirTag reveals how Amazon destroys rare books for AI training Vaslo · 38 pts · August 17, 2026 · 68% similar
- AirTag reveals Amazon is trashing rare books to train AI jefurii · 127 pts · August 17, 2026 · 67% similar
Discussion Highlights (20 comments)
ezfe
I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them. Instead, they enforce the copyright and force AI companies to shred books they want to ingest. edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
imperio59
Getting 10 million people to do anything is really, really hard. Getting 10 million people to spend hours scanning a book (which takes a really long time with a home scanner) sounds impossible :(
xvxvx
Pretty funny that they just took Anna’s archive and ingested it. As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
HedonicEscal8r
The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art. I support Anna's Archive, by the way. Information wants to be free.
shakna
I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves. Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book. So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
landgenoot
Isn't this a matter of regulation? I'm not sure about US, but in EU you have old houses/buildings that are protected. Sure, you can buy them, but you can't modify or destroy them (being cultural heritage).
mycall
Aren't AI companies all about the rare book auctions now?
whycombinetor
It's giving Vishnu, but the world cannot exist without Shiva.
c0lpan1c
that's ironic, the url annas-archive.gl is blocked by my local DNS category for AI Threat Detection.
tptacek
These stories are weird, because actual professional specialized book dealers pulp books by the millions . People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book. Even if they were literally burning these books to spite you, they'd be destroying an infinitesimal fraction of the books the book trade already destroys. It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed. The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure. But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
glimshe
The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022." Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
Cider9986
The AI companies should work with the Internet Archive to release the digitized copies once the copyright expires. Unrelated: So with this one copy BS are you not allowed to have backups of the data?
luciana1u
Someone should build the digital equivalent of a fire department. Train a model on the books, then if the originals get destroyed you still have the smoke.
BrenBarn
It's not a bad idea but we need a multi-pronged approach, with at least one other prong being "destroy the companies that are doing this".
WillAdams
"Whoever destroys a book destroys a link in the chain of human knowledge" -- Thos. Jefferson
ColdStream
The question I have is, do these companies keep copies of the scans after they have finished training on them? If so, then it isn't the worst outcome. Not great but at least the information is not completely destroyed forever just the original physical being of it. Deeper thought however, eventually this will all be lost to time and I suspect that about 99% of all printed materials probably would never be read again simply due to the huge volume of it and sheer obscurity. Ernest Becker and his work 'The Denial of Death' might have some thoughts on this. go to any second hand book store and just pick out something at random from the 1950's for instance, something about pottery or bird watching or whatever. The history of Bisbee Arizona, I don't know. Look up the author, see if they even left a trace of their work and the vast majority of the time they have already been forgotten to the great void of the universe. In the end, it all goes away. Clinging only creates pain. I'm not saying that we should let them just do this, I am just saying that long term it is a tough battle to fight only to lose the war.
SanjayMehta
Google Books was a great resource until the lawyers got involved. I was able to find and download (one screenshot at a time) a rare family history. The author died 100 years ago. The published disappeared 80 years ago. But now Google has locked it behind a limited preview. Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those systems.
altcognito
What evidence do we have that they are "destroying" books? I'm not saying this in their defense, but as someone who has worked at companies who has scanned books at scale, and generally speaking, I wasn't on site there, but I knew we/they were pretty delicate with the books. And while the kneejerk reaction might be "hey, why would they go through the effort?" -- my guess is that they are following or even hiring people that have done this process in the past (out of laziness) and just follow what works easiest. The literal machinery is not designed to destroy the books for various practical reasons. Books that are bound are easier to be kept in order and work with. Getting a flat scan is done with specialized tools, you don't need to put it on a plate (it would be too slow that way anyway) All of the above is just to justify my question: Who knows that the books are being destroyed? (I also agree with the general sentiment that there's a good chance these books are just cheap and bulk, they aren't pulling one of a kind rare books.)
mplewis
Can someone name a rare book that was destroyed as part of AI scanning? I want to know what kind of thing we're losing.
jupp0r
I highly doubt they destroy digital copies of the books after scanning. They will want to train their future models on the same content. So what prevents them from making these digital copies available to the public? Copyright!