It is a sign of the times that Amazon gets to call this fair use
sonicrocketman
105 points
79 comments
August 21, 2026
Related Discussions
Found 5 related stories in 44.9ms across 4,128 title embeddings via pgvector HNSW
- It's time Amazon played by the same rules as everyone else [video] whosgotch · 98 pts · August 11, 2026 · 55% similar
- The Amazon tax herbertl · 1089 pts · August 18, 2026 · 49% similar
- AirTag reveals Amazon is trashing rare books to train AI jefurii · 127 pts · August 17, 2026 · 48% similar
- Amazon, which started off selling books, is destroying rare texts to train AI rzk · 91 pts · August 17, 2026 · 47% similar
- Protester calls out Amazon CTO for allowing Israel to use their AI towards Gaza trymas · 27 pts · July 19, 2026 · 46% similar
Discussion Highlights (18 comments)
sonicrocketman
Original full title (too long): It is a sign of the times that Amazon gets to call this fair use while huge corporations try to sue the Internet Archive out of business.
tescreal
This has been reported of other major ai shops. It is a little chilling that they're going after rare books. For the record, i don't much have issue with llms and such. But... the wholesale destruction of printed books is a bridge too far.
Dylan16807
Copyright law prefers the destructive method. It's not that they're skirting the law but that the law is set up very badly in the first place.
WalterGR
The 404 Media report referenced in the article was submitted here: https://news.ycombinator.com/item?id=49330742 “We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility” (404media.co) 164 points | 3 days ago | 321 comments
IshKebab
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable. It does though. A lot of people seem to think "it's a book!!!" means it takes on some mystical intrinsic value. That's bullshit. There's an absolute mountain of worthless and near-worthless books. I bet the vast vast majority of these books are those books - after all that's what Amazon wants. A ton of human written text as easily as possible. They obviously aren't buying first editions of Oliver Twist, or even second editions of Harry Potter. > We’re not revealing the titles of the books included in the shipment we tracked Yeah... because then it would reveal how unimportant they are. This sort of outrage inflation is counter-productive. People see through it, lose trust, and then when something bad actually happens they won't believe you.
SturgeonsLaw
Getting real sick of capitalism systematically consuming and destroying everything it touches. These behaviours would be described as obsessive and pathological in any clinical setting, but when they're done for profit it's somehow normal.
like_any_other
> As mentioned elsewhere, this is also IP theft on a massive scale. Scanning a legitimately purchased book is IP theft? How can he hold such a copyright-maximalist view, and at the same time defend the Internet Archive?
anonymousiam
It's a win-win-win for Amazon if they destroy the source in the process of scanning it. They get the AI training data, and it's cheaper for them to destroy the book in the process. Destroying the book ensures that it will be more difficult for competitors to scan the same content, thereby increasing the value of the data they've scanned. Lastly, there's a misguided belief among some that it's somehow less of a copyright violation if the source is destroyed, but it's a violation either way unless the entity doing the scanning has permission from the copyright holder to make the copy.
8note
i dont get it. this is a complicated way of counting how many of each word is in the book. clearly fair use. does the author think cutting up a book is a copyright concern? theyre buying the books, its up to them what to do with their copy. if you want the books preserved, maybe fund your libraries to get a copy or two?
ChrisArchitect
Discussions: https://news.ycombinator.com/item?id=49330742 https://news.ycombinator.com/item?id=49310725 https://news.ycombinator.com/item?id=49068738
areoform
I think this is impressive backwards reasoning. Before my lifetime, copyright law all but choked and killed the public domain. And now everything is stale and the same. This particular battle was lost with Google Books and the attempt to make the world's largest library. The modern library of Alexandria. But then copyright lawyers got involved to get their pound of flesh. And here we are. I am upset about the destruction of knowledge. Paper is a superior storage medium to any hard-drive any day. We're recovering words from paper from over a thousand years ago. I think it's a mistake to not work with a non-profit, use cheap COTS non-destructive scanning, and write off the costs of rebinding them and rehousing them. Everyone is impressively short sighted.
Aurornis
The full headline tries to equate this to The Internet Archive’s recent legal troubles. For a refresher: The Internet Archive scanned books then shared them online, trying to claim that converting them into digital copies qualified as a derivative work. You don’t have to be a lawyer to see how that claim doesn’t hold up to any scrutiny. The LLM companies are not redistributing the works. They are using them for training. This blog post calls it IP theft, but it has actually been litigated in court already. Using books for training does not qualify as theft or redistribution, even though some people have different opinions about the moral angles. One of those same lawsuits also extracted a huge settlement from Anthropic for using digital downloads from pirate sites. The conclusion was that the only acceptable way for them to use the books is to buy them and scan them. The courts forced it to be this way. The current hand-wringing about the destruction of books is based on claims that they’re doing it to rare books that are also valuable. So far nobody has been able to actually provide an example of one of these books that is supposedly ultra-rare, but also valuable, but also only available at one of these book sellers that sells these things in bulk. We’re supposed to assume that one of these books might actually be super valuable but also super rare and also only available at these places they’re buying from.
wotamess
Everyone turning into Metallica. The whole "copyright infringement is theft" meme seems to have stuck. 20 years ago it was all "information wants to be free" and "infringement is not theft; being deprived of speculative profit is not a loss." How the turn tables.
schoen
It's not necessarily doctrinally that crazy in copyright law. Some courts have considered AI training on copyrighted works a "transformative" use (traditionally more protected in fair use analysis), while providing human beings access to the works a "consumptive" use (traditionally less protected). There's also a fair use consideration that favors noncommercial use compared to commercial use, but that's not the only question, and the statute doesn't say clearly how to combine the fair use factors. But it's possible under the Copyright Act that some commercial uses of copyrighted works could be considered fair uses while some noncommercial uses could simultaneously not be considered fair uses.
rkapsoro
I remember reading a scene where this exact thing was going on in Vernor Vinge's Rainbow's End: there, the books were being shredded and scanned as tiny bits and stitched together in software similar to shotgun genome sequencing. At any rate, I ought to pick it up again and finish reading it...
999900000999
I still think the solution would be for Amazon to release the EBook of any rare book they scan and destroy. Then setup a fund to compensate copyright holders. I guess publishers do have a right to pull certain works from print though.
satvikpendem
Well, this is what happens when they weren't allowed to use digital only copies, as they now have much more of an ability to get around copyright with physical items. The right thing would've been to allow digital items to also have the first sale doctrine, but we don't, so AI training companies must resort to destroying physical books in the process of scanning. The Cobra Effect strikes again.
tescreal
I wonder how much of AI is about owning all knowledge, deskilling the masses, reducing access, etc.? The end goal might be actual legit surfdom. Without the means to elevate the people's knowledge (and the companies have already stated they Benevolently Protect Users From Harmful Data™), it quickly becomes a situation where they control all narratives. It took less than a generation for social media to make the spectacular mess of discourse we now have. How long before all knowledge is mediated and moderated? Who's to say? I'll check back in 10 years :)