Live data from Hacker News

AI companies are shredding rare books

twitter.com

351–360 of 559 posts

Re: AI companies are shredding rare books

#352

Earlier quoted context omitted.

A friend proposed that copyright should just die with the author and / or their spouse and I'm left agreeing. I want books and music to be less strict on copyright. Some of my favorite YouTube channels break down music and songs, and go as far as recreating beats / tracks from famous hip hop songs, but someone at a record label company dings every one of their videos, they can barely sample a few seconds, its VERY CL…

20 years to make some money, and then we set the work free for the public benefit. If it's good enough for patents, I don't see why it isn't good enough for copyrights.

Either that or some exponentially increasing tax so that Disney can keep their vault. (I'm perfectly fine with them keeping it if they pay some proper taxes.)

Re: AI companies are shredding rare books

#353

Earlier quoted context omitted.

No, it isn't. Nothing about training a model requires this. You're thinking of copyright, which does treat information this way.

Sure, I'm conflating the tool w/ the people building the tool - but imo muddling the distinction is good, because the hard distinction is what tricks people into an unthinking mode.

On the contrary. The problem exists solely in copyright; AI companies are just users that happen to be popular to hate on, but this same legal situation applies to everyone else, including individuals. The issue under discussion is older than LLMs.

Re: AI companies are shredding rare books

#354
post #16

That's why archive.org should have never been sued for lending books they had physical copy of. This is the result. Publishers should be more careful what they wish for.

The archive.org story was more nuanced than that. If I recall correctly the full story was that they used to lend digital versions of books they physically bought and scanned with DRM to enforce a sort of one to one at a time restriction. But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently which started the debacle with the publishers.

If i recall correctly, they had permission from physical libraries to use their copies as well, so it wasn't just a single copy but many copies, just one of them was converted into digital. Still wasn't enough apparently...

Re: AI companies are shredding rare books

#355
post #340

Earlier quoted context omitted.

> for example, you wouldn't hire a 70 year old writer for your commercial project no matter how brilliant they are If you hire them, then you own the work you paid them to do, no?

Not in every country, and secondly if you're basing it on life of the author then that does't solve corporate copyright unless you tie it to the live of a particular employee. You could do "life of author or X years, whichever is longer". Or include a period after death. But you see how the complexities come in.

The complexities are the problem. I've always thought a fixed term is best. Then, you can purchase a work - it says Copyright on it. Then, you know that after dddd+term the copyright is lapsed. You don't need to hunt down the author to see if they died. No guessing, just written on the work that you purchased. No, don't have optional extensions - that just means you have to look it up. It should say it right there on the work you purchased when the copyright expires.

It is said that the vast vast majority of works don't earn anything significant after a few years in any case, meaning the only possible reason to have long copyrights is so that a very few people can get stinking rich. But those people already got rich, in the first few years.. society does not benefit from them getting richer.

20 years fixed term is my proposal.

Re: AI companies are shredding rare books

#356

This is potentially very bad. Like, Library of Alexandria or Council of Nicaea bad. We may never be able to recover the information if, say, one of these AI companies copied or translated it wrong then destroyed the source material. Maybe it's from bad OCR, or maybe from a bad actor - but there are a lot of ways history and information could change in this game-of-telephone like transfer of knowledge. What is the poi…

> What is the point of destroying the source material? I don't buy the copyright thing.

It is the copyright thing.

Despite what people say about scanning, the fact is, non-destructive scanning machines have been built and perfected long time ago. This was preferred in the past, back before some major kerfuffle with the publishers during COVID, but that incidentally happened to be before LLMs became a thing, so AI companies never had that option available.

Re: AI companies are shredding rare books

#357

Earlier quoted context omitted.

Yeah, I was thinking along these lines... Let's say one of the books to be digitized and destroyed is the sole remaining copy of a book from 1850, which is now considered public domain. On one hand, hoarding such a book, stealing its content from the public domain, locking its content behind a for-profit machine, and destroying the only remaining copy is clearly wrong. It's equivalent to stealing a public resource, j…

I think this misconstrus what public domain is. It provides a freedom to circulate, but not access to the material. It is not a public owned resource. Turning a copy over to the public or state might be an interesting requirement for obtaining a copyright, but instituting that fix for new works now would have a 70 year lag time. Think of it this way, if I copyright a book and put it in my dresser for 70 years, that d…

> It provides a freedom to circulate, but not access to the material. It is not a public owned resource

It increasingly looks like a grand compromise around copyright and AI is needed. Expanding public domain when it comes to AI companies in this respect seems merited. (The other bits are fair use if weights are opened.)

Re: AI companies are shredding rare books

#358

Earlier quoted context omitted.

> Why would DRM on a digital copy of a book matter? Because DRM is just a way to make "breaking copyright" more practically cumbersome. What's easier, breaking digital DRM for each and every E-book you find, or just establishing a single pipeline for scanning physical books?

breaking DRM is so easy my generation was doing it as kids, there is no technical obstacle there.

It's a manual process.

More importantly, it's also explicitly illegal. Destructive format-shifting is not. Thank copyright laws.

Re: AI companies are shredding rare books

#359
post #16

That's why archive.org should have never been sued for lending books they had physical copy of. This is the result. Publishers should be more careful what they wish for.

Publishers don't care if rare books get shredded?

There are many reasons but it's just rare books - in general: they are setting a precedent for people to pirate instead of using libraries. In general companies are trying really hard to make people switch back to torrenting, pirate sites, sharing media etc.

Re: AI companies are shredding rare books

#360
post #289

Earlier quoted context omitted.

This entity is destroying books and make them only accessible via the "intelligence" exposed by their LLMs. So I ll give you two options... * their values are aligned with the best interests of humanity * their values are not aligned with the best interests of humanity. Now go ahead and read my mind!

> make them only accessible But that's a lie. They're not destroying the last copy of books.

TFA says so...
Post reply on HN