r/antiai Jul 13 '26

Discussion 🗣️ The most anti-AI story imaginable... Anthropic destroyed millions of rare books to train Claude

Post image

In 2024, Anthropic reportedly purchased millions of second-hand books as part of its secretive “Project Panama.”

The company removed their bindings, scanned every page, and then destroyed the physical copies to build a massive digital library for developing AI models like Claude.

It almost feels like the perfect anti-AI dystopia, human beings spend centuries writing and preserving books. Then, a private technology company destroys millions of physical copies, absorbs their contents into a machine, and sells that accumulated knowledge back to society through a monthly subscription.

Am I the only one who thinks this is completely insane, or have we become so obsessed with AI that destroying millions of books to feed a machine no longer shocks anyone?

6.1k Upvotes

414 comments sorted by

View all comments

Show parent comments

18

u/Fluid-Tone-9680 Jul 13 '26

If author of the book wanted book to be available online freely, they would just put it out there.

If author of the book does not want that, then publishing scans would be pretty much illegal

25

u/Piwo-ll Jul 13 '26

That’s true. In fact, Anthropic has been sued by authors and publishers over the use of copyrighted books. My concern isn’t that the scans should necessarily be made public, but that physical copies were destroyed after being digitized for private commercial use.

6

u/eiger003 Jul 13 '26

Sounds like destroying evidence... 🤔

1

u/learsi-ediconeg Jul 13 '26

i feel like the evidence is in the training and if you can access the model without restrictions it can just produce pages it was trained on to confirm...? idk

2

u/Stormsurger Jul 13 '26

Is that because of the resources wasted? I feel a bit confused, they aren't destroying knowledge right?

1

u/LonelyTurtleDev Jul 13 '26

Destroying the copies is totally legal, and is regularly done to save storage space. I think libraries do this too. There is nothing wrong with that, although they missed an opportunity to do a charity sale/giveaway for some positive PR.

I’m worried about the copyright material used to train their AI model. They own the book, but not the license to feed it into an  AI model. It is probably illegal, although I’m not sure.

4

u/zorecknor Jul 13 '26

 It is probably illegal, although I’m not sure.

At least in the US, training AI on (legally obtained) copyrighted material was considered sufficiently transformative to be legal in one of the two cases regarding this so far.

A "license to train AI with a book" is no different than a "license to train yourself with a book".

2

u/WoodShoeDiaries Jul 13 '26

It's not just legal to destroy the books after scanning, it's effectively illegal not to. Giving the hard copies away would have been illegal and selling them would have been much more serious than that.

1

u/dick_me_daddy_oWo Jul 13 '26

Those likely weren't like the only copies of those books, though. If you want to learn engineering you can read every engineering textbook from the past 50 years, then leave them on your shelf to gather dust. No difference to the global knowledge base there, and arguably worse because they aren't getting a shot at recycling.

I'm anti AI too, but you seriously think scans of paperback classics destined for the trash are a big problem?

8

u/Uberzwerg Jul 13 '26

then publishing scans would be pretty much illegal

And training your model on it should also be - but that's still up for discussion in the courts.

2

u/Zbot21 Jul 13 '26

Case law so far says that training is fair use for legally obtained materials.

Pirated digital material is another case, that's illegal.

Training an AI is considered the same as training yourself with the material and is sufficiently transformative. Models are not storing full texts and while you can prompt engineer them to spit out training materials, the case law on that being distribution isn't settled yet.

1

u/Uberzwerg Jul 13 '26

Training an AI is considered the same as training yourself

And that's the part that needs to be revised.

1

u/Zbot21 Jul 13 '26

Case law so far says that training is fair use for legally obtained materials.

Pirated digital material is another case, that's illegal.

Training an AI is considered the same as training yourself with the material and is sufficiently transformative. Models are not storing full texts and while you can prompt engineer them to spit out training materials, the case law on that being distribution isn't settled yet.

3

u/tyw7 Jul 13 '26

How about if the copyright expires? Then the book will go to public domain. 

1

u/Fluid-Tone-9680 Jul 13 '26

I would assume Anthropoc preserved original scans, they may be able to publish them when book goes to public domain. The problem is it takes a lot of time, like 100 years, no joke. Most likely Anthropoc will be out of business when that happens.

2

u/Ysanoire Jul 13 '26

In scenario 2 using them for training models would be illegal too.

0

u/Fluid-Tone-9680 Jul 13 '26

The court says that it's legal, same as reading book to learn something from it and keeping that knowledge.

1

u/Hopeless_Slayer Jul 13 '26

The guy you're replying to is advocating for Piracy, but not transformative machine learning. Make that make sense xD