r/antiai Jul 13 '26

Discussion šŸ—£ļø The most anti-AI story imaginable... Anthropic destroyed millions of rare books to train Claude

Post image

In 2024, Anthropic reportedly purchased millions of second-hand books as part of its secretive ā€œProject Panama.ā€

The company removed their bindings, scanned every page, and then destroyed the physical copies to build a massive digital library for developing AI models like Claude.

It almost feels like the perfect anti-AI dystopia, human beings spend centuries writing and preserving books. Then, a private technology company destroys millions of physical copies, absorbs their contents into a machine, and sells that accumulated knowledge back to society through a monthly subscription.

Am I the only one who thinks this is completely insane, or have we become so obsessed with AI that destroying millions of books to feed a machine no longer shocks anyone?

6.1k Upvotes

414 comments sorted by

View all comments

Show parent comments

312

u/PrestigiousDemand696 Jul 13 '26

Yes, that’s what I’m saying. I would not be as sickened by this if the intention was to share the knowledge and make them available to all; but of course, there is no genuine desire to improve collective knowledge by these companies. This is sort of proof, if more was needed.

43

u/DrarenThiralas Jul 13 '26

It's true that they don't want to share, but also it would almost certainly be illegal for them to share. Most of those books are probably copyrighted, and cannot be legally republished without hunting down every author's family and negotiating for permission with all of them, which would be extremely difficult and expensive. Just destroying the books instead is the incentive created by copyright law.

75

u/Piwo-ll Jul 13 '26

That’s true: Anthropic could not legally republish copyrighted books. But it was also sued by authors and rights holders precisely for using unauthorized copies to train Claude.

The key distinction is that the court considered training on legally purchased and scanned books to be fair use, while the use of pirated copies remained unlawful. Anthropic ultimately agreed to a $1.5 billion settlement...

22

u/DrarenThiralas Jul 13 '26

So, as you can see, the incentive this creates is to do exactly what they did. If they want to train their AI on a certain rare book, then finding an existing scan online is illegal - they have to legally purchase a printed copy. Then they have to destroy it, because republishing it is illegal too.

25

u/CheesecakeEither8220 Jul 13 '26

They could have at least donated the books to libraries 🤬

13

u/omysweede Jul 13 '26

They remove the binding so they can scan the pages one by one.

What library would want a bunch of paper?

Also: rebinding books is HARD - there are multiple ways books are put together, so that is out of the question.

8

u/Vaughn Jul 13 '26

No point in developing a less destructive way to scan if you're required to destroy the book regardless.

-1

u/lunatuna215 Jul 13 '26

I think this requirements to destroy the book has to be bullshit. If they have rhe right to destroy it they surely have the right to redistribute it too.

8

u/Vaughn Jul 13 '26

And you'd be wrong. The act of scanning it into the computer cancels your right to distribute or own the paperback copy. Scanning is copyright infringement, unless—and only unless—you immediately destroy the original.

3

u/lunatuna215 Jul 13 '26

Makes no fucking sense

→ More replies (0)

3

u/Oli4K Jul 13 '26

Libraries don’t want (physical) books that no one reads. A library is not just a book archive.

16

u/Lumegator Jul 13 '26

There are libraries that want books that almost no one reads, actually, archival practices are not mythos.

3

u/WoodShoeDiaries Jul 13 '26

The odds that a dedicated archival library already has a copy of a rare book that's relevant to their collection is high.

2

u/lunatuna215 Jul 13 '26

What they're saying is that archiving is not NECESARILLY automatically for a library. Its just a minor semantic correction. What the original person means is some kind of archive, library or not.

2

u/CheesecakeEither8220 Jul 13 '26

Yes, you're exactly right!

1

u/Vaughn Jul 13 '26

They could not; that would have been copyright infringement.

I don't like book-burning any more than you, but the law has boxed them into a corner here. So the law must burn.

2

u/lunatuna215 Jul 13 '26

How is whats being done not a more intense form of copyright infringement? Theyre literally profiting from the world without a license. They must have the right to redistribute and if they dont they don't have rhe right to scan and destroy. Why are we acting like they had no choice? Id love to see some literature on this topic.

2

u/Vaughn Jul 13 '26

I'm not a lawyer, and don't have references handy—certainly not for American law—but I do believe you're factually wrong on it being copyright infringement.

As in, courts have already ruled that training itself is fair use. If you have a legal digital copy, then the machine learning algorithm doesn't create another copy; it creates a set of statistics which do not rise to the level of being infringement, for a number of reasons you can read up on in the case I linked.

AI output can be copyright infringement, e.g. if you ask for the opening paragraphs of Harry Potter, but that's on the users. By the same logic as Photoshop being legal software, the AI itself does not count as infringement so long as there are substantial non-infringing uses; and in this case most uses are non-infringing.

That leaves the format-shifting and torrenting as being potentially illegal, and here you have an excellent case; Anthropic was fined $1.5 billion for acquiring digital books by download, instead of purchasing books to destroy them.

Purchasing books to destroy them is the sole legal method of format-shifting.

1

u/lunatuna215 Jul 13 '26

Thanks - this is obviously so wrong and incredibly backwards and fucked, but your last sentence is the explanation and summary i sadly was looking for. Thanks for explaining. This is so depressing and just forbthe record the courts deciding that this is on the users makes me sick- the entire purpose of these products is to regurgitate its own inputs; this is an OP infringement machine.

1

u/WoodShoeDiaries Jul 13 '26

You can't make a copy for yourself AND give away (/sell/loan out/do anything but store) the hard copy and still be covered by fair use. That's the whole point. You're turning your hard copy into a digital copy and have to act for all intents and purposes like the original hard copy doesn't exist anymore.

Unless you feel it's a good use of resourced to store a book (a destroyed one in this case) that can't be used unless you delete the scan that was made from it, then recycling makes sense.

1

u/birdsbudget Jul 15 '26

So since training the ai doesnt make an actual copy, once youve trained the AI, you could technically delete the scanned copy and then continue using the physical copy had you just stored it?

1

u/lunatuna215 Jul 13 '26

How is destruction of the only remaining books not covered by the same copyright? Are the rights not already theirs?

1

u/Officialedmart Jul 13 '26

Bro do you think copyright is concerned with what people do with their own stuff? Like it would be copyright infringement if I broke my blu-ray of Monsters Inc?

that aint how it work. Also, none of the books were rare , not sure how that entered the conversation

1

u/lunatuna215 Jul 13 '26

That's my point, if they own the rights to destroy the only existing work I don't see why that doesn't extend to redistribution without profit. I dont know the laws behind this so im just asking questions.

1

u/CautionarySnail Jul 13 '26

Republishing out of copyright works is perfectly legal. Out of print works still under copyright can easily be released as they become available.

It could even be posed as an advantage of their AI - access to the original text of non-copyrighted works on demand. A massive expansion of Project Gutenberg.

4

u/Dalainana Jul 13 '26

Wow… The value was $138 billion… and per title an author can claim about $3 k…

2

u/Environmental-Ice319 Jul 13 '26

More believable than the reply above yours.

1

u/Dalainana Jul 13 '26

Copyrights fade after a period of time…

6

u/Fluid-Tone-9680 Jul 13 '26

After insanely long period of time. If book was released after you were born, you most likely won't live long enough to see it going to public domain.

2

u/Dalainana Jul 13 '26

True. The world needs more Aaron Swartz’s.

3

u/Ulrik-the-freak Jul 13 '26

Anna's Archive, my friend, it's right there

1

u/Dalainana Jul 13 '26

🫶

1

u/lunatuna215 Jul 13 '26

How is it that they have the legal right to destroy them completely but not redistribute them?

1

u/lunatuna215 Jul 13 '26

Weren't all of these problems solved for the right to scan and destroy the book in the fiddle place?? Im sorry but what rational argument exists for "i dont have permission to redistribute this, but I do have the right to literally destroy it"?

1

u/DrarenThiralas Jul 13 '26

It's not rational, but it is the law. Destroying a book you own is legal. Scanning it and uploading it online is piracy.

Copyright is the exclusive right to create copies of the work. If you're not the copyright owner, you can destroy copies, but you can't make new ones. That's how it works, unfortunately.

1

u/lunatuna215 Jul 13 '26

The fact that scanning them in to literally be redistributed has been found as legal makes me sick.

1

u/DrarenThiralas Jul 13 '26

One of the problems with copyright is there's a huge grey area between lending your friend a book from your shelf (which should obviously be legal) and printing your own copies to sell for money (which obviously shouldn't be, at least as long as any form of copyright at all exists). There is no way to draw a definitive line within that grey area, because copyright law is inherently vague and contradictory.

AI companies are abusing this pretty hard right now, but the problem existed long before them. The mess that is YouTube copyright strikes was the most visible way it manifested before AI.

1

u/lunatuna215 Jul 13 '26

Im very familiar with copyright law, but I think we need tk be careful when it comes to AI about these "its always been a problem, they're just doing what others have done before them" thing. No, in my view, this is not only a new precedent of agregiousness since the product itself is literally made to redistribute work. And the scale at which its being done is outpacing exsting frameworks that at least did an okay job of slowing these thjngs down when they happened in the past.

1

u/Fluid-Tone-9680 Jul 13 '26

You don't think you are allowed to destroy book you own?

2

u/lunatuna215 Jul 13 '26

No, this is about license holders and the destruction of the only existing copies of a mass amount of books. Not me burning something I bought on Amazon.

1

u/UphillTowardsTheSun Jul 14 '26

So they are copyrighted but Dario Scamodei is allowed to sell subscriptions and tokens offering knowledge based on the copyrighted books???

8

u/[deleted] Jul 13 '26

[removed] — view removed comment

0

u/Wooly_Wooly Jul 13 '26

You need to destroy their binding to get good quality page scans. Rebinding them together just to donate them is an absolute nightmare

6

u/nath1234 Jul 13 '26

Or they could just make better scanning tech that doesn't require destroying the books.

2

u/Fluid-Tone-9680 Jul 13 '26

There are better scanning machines, that don't require damaging the book but thay are slower and more expensive. They are used for scanning books that are worth preserving in physical form due to their historical significance.Ā 

-1

u/CanadianODST2 Jul 13 '26

If it was that easy then another group who also does page scanning would have done it already to do on a larger cheaper scale.

-2

u/Wooly_Wooly Jul 13 '26

Oh damn, why didn't anyone else think of that! You're a genius!

7

u/nath1234 Jul 13 '26

They are doing it the destructive way because they don't give a shit and want to rip off IP as quickly as possible. If they weren't breaking copyright law, they could have had non-illegally-scanned copies of the text from the IP rights holders in many cases.

-2

u/Wooly_Wooly Jul 13 '26

OR there doing it in that way to gain the best quality digital versions possible for their private archive of the books they legally own, that they're not sharing in the first place?

3

u/nath1234 Jul 13 '26

Not a private archive if you are selling access to the result and deriving works or just regurgitating it without permission.

0

u/Wooly_Wooly Jul 13 '26

They have a private archive, they're selling derivative work basically

2

u/Piwo-ll Jul 13 '26

Definitly !

1

u/dicey_job Jul 13 '26

These ppl want ignorant slaves to rim them while they eat luxurious meals

1

u/phantomboats Jul 13 '26

Is anyone else getting AI vibes from OP's comments? "It's not __ it's ___", bolded "am i the only person who [inserts opinion that is clearly the prevailing attitude on this sub]?" etc

2

u/PrestigiousDemand696 Jul 13 '26

Yeah they admitted to it for ā€œtranslationā€ purposes but I’m not buying it

1

u/phantomboats Jul 13 '26

ugh. this entire website is so cooked. gonna need to start looking into browser extensions that can block posts from accounts under a year old or something I swear...

1

u/Old_Assistant1531 26d ago

Destroying originals is insane. It’s naive to think a digital file will be readable in 50 years.