AI companies are shredding rare books

anon373839 757 points 479 comments July 27, 2026
twitter.com · View on Hacker News

https://xcancel.com/HedgieMarkets/status/2081534588485296565

Discussion Highlights (20 comments)

sherr

I see mentions of Bradbury's "Fahrenheit 451" in that thread but what this really seems to be mostly like is Vernor Vinge's "shred and scan" factory in his novel "Rainbows End".

thechao

Which book that was rare was destroyed? I'm interested to know a few titles.

est31

> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate. Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright? IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies. Scanning books you own should be legal from a copyright point of view, and not require shredding. Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work.

stuartjohnson12

I think people are, on the whole, too precious about old things. In the case of books produced after major commercial printing began, I don't believe it is the paper that imbues the book with historical value. Indeed, I think there's a high chance that this process increases preservation of the most relevant part of the media - the actual content! There are tons of old rolls of film slowly rotting away in warehouses that were never digitised. Even for beloved media, the BBC occasionally tracks down an old lost episode of Dr Who. For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.

ACCount37

The publishers sued AI companies for training on shadow library data, hoping to negotiate content deals for big $$$ down the line. Instead, they got analog hole'd. Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers. What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after the books that there's still copyright on. Second, when it comes to books, "old" doesn't mean "valuable" - plenty of libraries destroy old books because there's no demand for them, and storage costs you. This is how those scanning companies get books for so cheap.

vessenes

I don't see any proof of shredding here. Most book scanners I'm aware of are from Google's scanning days, and those had cameras plus page turning. If we use 'shredding' to mean A book is laid flat, its cover is removed, and then a paper cutter cuts through the binding to create a flat stack of sheets, which are then fed to a sheet feeder, then I could maybe imagine this is better than a page-turning scanner. But, sheet feeding old paper sucks shit, bro. It's not fun. Upshot, I think we'd like to hear from an anonymous frontier lab employee here to see what's going on -- there are a lot of books in Anna's archive available at considerably less difficulty.

lousken

That's why archive.org should have never been sued for lending books they had physical copy of. This is the result. Publishers should be more careful what they wish for.

Good4boothee

> A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time. Really? That sure wasn't a thing when one startup got sued for streaming from its wall od dvd-players, and they adhered to 1 disc = maximum 1 stream at same time.

retinaros

if they do this you can foresee what else they can do

timcobb

This kind of reads like a blood libel. My guess is they're buying all those books that University libraries are throwing away these days (to see other HM threads for that), kinda sad but probably aren't "rare books" in the way people are thinking

azan_

The comments there are absolutely unhinged. There are some good reasons for being anti-AI, but why dilute it with this kind of bullshit: > It is equivalent to book burning in the past. A form of thought control

johnxianren

I have zero proof for this, but just a what if: what if Anthropic's strict anti-China stance actually means the Chinese training corpus is way more valuable than people realize?

skybrian

It’s unclear whether ISBNdb will scan books without ISBN’s, which were invented in the late 1960’s. Customers appear to be ordering books to be scanned by ISBN? Here is one book seller’s experience: > Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases. Article is paywalled, but I saved a few quotes here: https://skybrian-links.exe.xyz/post/1026

hackernudes

Also discussed here https://news.ycombinator.com/item?id=44381838 from June 2025.

enaaem

This is the destruction of Western civilisation.

Springtime

It seems a key contention of theirs is the possibility that rare books are being destroyed this way, yet the things they cite don't seem to suggest this (based on their paraphrasing), they just throw the following at the end to make it seem like it's occurring to irreplaceable books: > You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. Is there evidence of this? Since otherwise they could very well be describing what is only occurring to in-print or non-rare books. (This is a genuine question since their post doesn't shed any light on it.)

pu_pe

From what I understand, the rare books in question are not some historically relevant medieval manuscripts, but rather some relatively recent books (still under copyright) for which there are few print copies available for purchase. I am not sure that physically destructing one copy of this type of book to preserve its contents digitally is so bad. Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.

musha68k

At the very least why not upload the scanned books to the internet archive while already at it? This is highly disturbing news; is this standard practice? What did Google Books do before?

qsera

I will make a robot scanner for books. I will then scan all the books in my state libraries and make a digital copy of them (without destroying them) before these things come for them. I wish...

D13Fd

This is the result of our copyright law in the United States, which is extremely tilted to favor authors and publishers. The judge made exactly the right call and the companies are following the law. The fix here is to change the law to permit training AI without destroying the original materials. But that is going to be a heavy lift.

Semantic search powered by Rivestack pgvector
15,062 stories · 140,779 chunks indexed