Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

91–100 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#91

To clarify: Are they scanning and destroying a single copy of Book X or are they buying up all copies of book X, scanning it once, then destroying all copies of book X they can get their hand on?

They are ordering books with ISBN. So I take that they are tracking what books they have scanned or pirated already and only picking up what they are missing. As just ordering mass bulk and getting 20 of the same encyclopaedia would be waste.

And I guess something like encyclopaedia would be good example of book they scan. At one point popular, but with most copies destroyed as no one actually wants them anymore.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#92
post #6
post #2

I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them. Instead, they enforce the copyright and force AI companies to shred books they want to ingest. edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than…

It is also the case that the copyright holders are often putting restrictions around use of electronic forms that are driving the desire to use physical copies. I doubt AI companies would use a single physical book if they could avoid it - absent the legal cloud over electronic rights. I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to for…

> I doubt AI companies would use a single physical book if they could avoid it

They just don't want to pay what the copyright holders want to charge

Re: AI companies destroy physical books – let's scan rare books before it's too late

#94
post #77

Earlier quoted context omitted.

So, have you tried finding out what the programming was in October 1994? Or what cultural ephemera appeared in the TV guides of that era alongside the schedules? Either there's a copy for the week you want in an archive, or somebody's got one for sale, or most often neither. This can piss you off, if as it happened you had a reason to care.

That would make sense as an argument if the natural endpoint of these copies was preservation, and AI was disrupting that. But the natural next step for virtually all these books is to be recycled, not preserved. Books are generally not preserved . It is extremely normal for them to be pulped. Millions and millions of books are pulped every year.

[deleted]

Re: AI companies destroy physical books – let's scan rare books before it's too late

#95

Earlier quoted context omitted.

They don't "force" anything. Trillion dollar AI companies and their owners have as much agency as book publishers.

They are "forced" to do this because that's what they have to do to abide by copyright law. They can't create a digital duplicate without destroying the original.

If that's true, then why did they pirate so many books?

Re: AI companies destroy physical books – let's scan rare books before it's too late

#96

The AI companies should work with the Internet Archive to release the digitized copies once the copyright expires. Unrelated: So with this one copy BS are you not allowed to have backups of the data?

I entirely believe the litigation brought against Internet Archive was secretly sponsored by these exact organizations, because they want to monopolize information to train models. No data => No models => No competition.

IA was in hot water already even before ChatGPT came out.

> ChatGPT […] originally released on November 30, 2022

https://en.wikipedia.org/wiki/ChatGPT

> On March 24, 2020, following shutdowns caused by the COVID-19 pandemic, the Internet Archive opened the National Emergency Library, removing the waitlists used in Open Library and expanding access to these books for all readers. More than one user could borrow a book at the same time. Two months later, on June 1, the National Emergency Library (NEL) was met with a lawsuit from four book publishers. Two weeks after that, on June 16, the Internet Archive closed the NEL, and the prior Open Library CDL system resumed after the 12 weeks of NEL usage.

https://en.wikipedia.org/wiki/Hachette_v._Internet_Archive

Re: AI companies destroy physical books – let's scan rare books before it's too late

#98
post #13

I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves. Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book. So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them…

Since when do AI companies care about the law? Most of their training data is pirated.

> Since when do AI companies care about the law? Most of their training data is pirated.

I imagine since the law recently cost one of them truckloads of money for their violations of it?

Re: AI companies destroy physical books – let's scan rare books before it's too late

#100
post #86

This whole situation is such a disgusting consequence of copyright law. The most frustrating part is that its so artificial. It is 100% the consequence of stupid laws.

> It is 100% the consequence of stupid laws.

More like the consequence of being unwilling to change stupid laws once the stupidity of them is discovered. Nope. Gotta double down on the stupidity instead...

Post reply on HN