Live data from Hacker News

Penguin Random House underscores copyright protection in AI rebuff

thebookseller.com

21–30 of 57 posts

Re: Penguin Random House underscores copyright protection in AI rebuff

#21

It's hard to understand the position of AI companies, when their AIs can be prompted to reproduce copyrighted works verbatim, e.g. https://imgur.com/a/XBO2B7V You can argue that a game that old ought not to be still in copyright, and I'd support that position. But it is in copyright, and I'd rather a world where we curtail copyright terms but enforce them fairly rather than a world in which we stifle culture by autom…

It's a complicated mess.

Reason being: I, too, am physically capable of reproducing copyrighted works verbatim when correctly prompted. Or at least, can do so as well as current GenAI can in certain specific cases, and know how to use a photocopier in others.

Is the copyright violation the user asking for this, or is it the model having the capability?

If it's the mere capability, the inventors of copy-paste and the camera and voice recorder apps have a problem, along with all search engines that ever index pirated material.

If it's the users, given that copyright databases are much too large for any human to actually know exhaustively what is and isn't in them, there's a real danger that normal people will accidentally do exactly this.

Re: Penguin Random House underscores copyright protection in AI rebuff

#22
post #21

It's hard to understand the position of AI companies, when their AIs can be prompted to reproduce copyrighted works verbatim, e.g. https://imgur.com/a/XBO2B7V You can argue that a game that old ought not to be still in copyright, and I'd support that position. But it is in copyright, and I'd rather a world where we curtail copyright terms but enforce them fairly rather than a world in which we stifle culture by autom…

It's a complicated mess. Reason being: I, too, am physically capable of reproducing copyrighted works verbatim when correctly prompted. Or at least, can do so as well as current GenAI can in certain specific cases, and know how to use a photocopier in others. Is the copyright violation the user asking for this, or is it the model having the capability? If it's the mere capability, the inventors of copy-paste and the…

Ridiculous. The copyright violation is in its use for training data. And then its doubled down by the user asking for copyrighted material and getting it verbatim.

It's not complicated. It's only complicated because it might be in the way of some people making billions or trillions.

Let's stop with the analogies. Analogies won't get you anywhere.

- if you recorded a concert with a tape recorder, then distributed it online for money, that was a copyright violation.

- if you recorded a CD with a tape recorder, then sold copies, that was a copyright violation.

- If you used a camera to film a movie being shown in a theatre, then distributed it online, that is a copyright violation.

Stop the BS.

Re: Penguin Random House underscores copyright protection in AI rebuff

#23

Which country has the most favorable 'fair use' laws, and why wouldn't big companies train their models there?

IIRC Japan has had at least one court ruling that training on copyrighted data is fair use (or a version thereof).

Japan does not have US-style fair use, their copyright exemptions are closer to UK/EU-style fair dealing (with stricter enforcement and a shorter copyright term on work-for-hire, though).

Re: Penguin Random House underscores copyright protection in AI rebuff

#24
post #21

Earlier quoted context omitted.

It's a complicated mess. Reason being: I, too, am physically capable of reproducing copyrighted works verbatim when correctly prompted. Or at least, can do so as well as current GenAI can in certain specific cases, and know how to use a photocopier in others. Is the copyright violation the user asking for this, or is it the model having the capability? If it's the mere capability, the inventors of copy-paste and the…

Ridiculous. The copyright violation is in its use for training data. And then its doubled down by the user asking for copyrighted material and getting it verbatim. It's not complicated. It's only complicated because it might be in the way of some people making billions or trillions. Let's stop with the analogies. Analogies won't get you anywhere. - if you recorded a concert with a tape recorder, then distributed it o…

> The copyright violation is in its use for training data.

Every textbook, every educational TV show or YouTube video, every artist whose works have been shown to me in my school lessons in art, music, graphic design, usw. have all been copyrighted.

With one exception: Shakespeare.

Even the bible, being a translation*, had copyright notices on it.

But worse than that: if it were so even for scraping the whole web and training an AI on it… what do you think Google is? Page Rank is a big ol' matrix multiplication, and it spits out quotes verbatim.

> It's not complicated. It's only complicated as it might be in the way of some people making billions.

You know some of these LLMs can be downloaded and run locally, right? For free, even.

"Not making money" is absolutely not sufficient to remove a claim of copyright infringement; and conversely these models are too useful to be simply ignored — they're now in the realm of national strategic thinking.

> Let's stop with the analogies. Analogies won't get you anywhere.

They are the only thing we have, as this didn't exist before, and we need to create a new legal framework for them.

Even copyright itself as a legal idea comes from analogies starting at least as far back as a 6th century Irish dispute where King Diarmait Mac Cerbhaill gave the judgement "To every cow belongs her calf, therefore to every book belongs its copy."**

> Also, if you recorded a concert with a tape recorder, then distributed it online for money, pretty sure that was a copyright violation.

Yes, that's the point.

And yet I am not forbidden from being online due to the fact that I have the capability to do so.

Only the actual performance of this is an offence, not the mere capability.

* All translation is necessarily also interpretation, specifically this one: https://en.wikipedia.org/wiki/Good_News_Bible

** Or at least that's the modern English translation of what he said, presumably either in in pre-medieval Latin or Old Irish.

Re: Penguin Random House underscores copyright protection in AI rebuff

#25
post #21

Earlier quoted context omitted.

It's a complicated mess. Reason being: I, too, am physically capable of reproducing copyrighted works verbatim when correctly prompted. Or at least, can do so as well as current GenAI can in certain specific cases, and know how to use a photocopier in others. Is the copyright violation the user asking for this, or is it the model having the capability? If it's the mere capability, the inventors of copy-paste and the…

Ridiculous. The copyright violation is in its use for training data. And then its doubled down by the user asking for copyrighted material and getting it verbatim. It's not complicated. It's only complicated because it might be in the way of some people making billions or trillions. Let's stop with the analogies. Analogies won't get you anywhere. - if you recorded a concert with a tape recorder, then distributed it o…

Well, copyright law has literally everything to do with reproducing content and has nothing to do with consumption of content.

So I have to disagree. Copyright doesn’t stipulate the terms of your consumption of media. Sure, people can write licenses and whatnot, but that’s not copyright, that’s a license (for example, the TOS of NYT website may dictate your rights to scrape it)

Copyright law is woefully under-prepared to deal with the challenges of llms. If someone has a photographic memory, it’s not illegal for them to read the book, it’s illegal for them to reproduce it. That’s essentially what we’re seeing with LLMs.

In all of your examples, the illegal part is really the distribution so far as copyright is concerned.

Re: Penguin Random House underscores copyright protection in AI rebuff

#26

Earlier quoted context omitted.

Ridiculous. The copyright violation is in its use for training data. And then its doubled down by the user asking for copyrighted material and getting it verbatim. It's not complicated. It's only complicated because it might be in the way of some people making billions or trillions. Let's stop with the analogies. Analogies won't get you anywhere. - if you recorded a concert with a tape recorder, then distributed it o…

Well, copyright law has literally everything to do with reproducing content and has nothing to do with consumption of content. So I have to disagree. Copyright doesn’t stipulate the terms of your consumption of media. Sure, people can write licenses and whatnot, but that’s not copyright, that’s a license (for example, the TOS of NYT website may dictate your rights to scrape it) Copyright law is woefully under-prepare…

So how do you suggest LLMs could prevent reproduction of copyrighted material, if they're doing so now.

Re: Penguin Random House underscores copyright protection in AI rebuff

#27
post #24

Earlier quoted context omitted.

Ridiculous. The copyright violation is in its use for training data. And then its doubled down by the user asking for copyrighted material and getting it verbatim. It's not complicated. It's only complicated because it might be in the way of some people making billions or trillions. Let's stop with the analogies. Analogies won't get you anywhere. - if you recorded a concert with a tape recorder, then distributed it o…

> The copyright violation is in its use for training data. Every textbook, every educational TV show or YouTube video, every artist whose works have been shown to me in my school lessons in art, music, graphic design, usw. have all been copyrighted. With one exception: Shakespeare. Even the bible, being a translation*, had copyright notices on it. But worse than that: if it were so even for scraping the whole web and…

I agree that copyright law is woefully unprepared, and that LLMs have become useful enough it will be difficult if not impossible to stop their development.

> Page Rank is a big ol' matrix multiplication, and it spits out quotes verbatim.

That makes no sense. Google Books was blocked exactly because it was reproducing enough material to be infringing. Google search usually does not copy enough material that it might be considered infringing (even though I think news sites have sued over this in some jurisdictions, right?)

> And yet I am not forbidden from being online due to the fact that I have the capability to do so.

It's fairly obvious that the people running LLMs don't want to prevent them from infringing copyright, particularly in image gen.

Re: Penguin Random House underscores copyright protection in AI rebuff

#28

Earlier quoted context omitted.

Well, copyright law has literally everything to do with reproducing content and has nothing to do with consumption of content. So I have to disagree. Copyright doesn’t stipulate the terms of your consumption of media. Sure, people can write licenses and whatnot, but that’s not copyright, that’s a license (for example, the TOS of NYT website may dictate your rights to scrape it) Copyright law is woefully under-prepare…

So how do you suggest LLMs could prevent reproduction of copyrighted material, if they're doing so now.

I think that’s a more interesting question. I’m not sure how to do it, but I think that finding a way to stop the reproduction of copyrighted content is probably the missing piece.

If there was a monetary penalty for reproduction of copyright works (apply human laws to machine), then I bet these companies would quickly figure out how to fingerprint output and match it to source data before sending it to the user.

Re: Penguin Random House underscores copyright protection in AI rebuff

#29

Earlier quoted context omitted.

So how do you suggest LLMs could prevent reproduction of copyrighted material, if they're doing so now.

I think that’s a more interesting question. I’m not sure how to do it, but I think that finding a way to stop the reproduction of copyrighted content is probably the missing piece. If there was a monetary penalty for reproduction of copyright works (apply human laws to machine), then I bet these companies would quickly figure out how to fingerprint output and match it to source data before sending it to the user.

So I go into a cinema with my recording equipment, record the movie, step outside and start selling downloads, and when the cops come, I say:

"I’m not sure how to do it, but I think that finding a way to stop the reproduction of copyrighted content is probably the missing piece."

Really?

Re: Penguin Random House underscores copyright protection in AI rebuff

#30
post #24

Earlier quoted context omitted.

> The copyright violation is in its use for training data. Every textbook, every educational TV show or YouTube video, every artist whose works have been shown to me in my school lessons in art, music, graphic design, usw. have all been copyrighted. With one exception: Shakespeare. Even the bible, being a translation*, had copyright notices on it. But worse than that: if it were so even for scraping the whole web and…

I agree that copyright law is woefully unprepared, and that LLMs have become useful enough it will be difficult if not impossible to stop their development. > Page Rank is a big ol' matrix multiplication, and it spits out quotes verbatim. That makes no sense. Google Books was blocked exactly because it was reproducing enough material to be infringing. Google search usually does not copy enough material that it might…

I disagree: I think copyright law is wholly prepared.

If some thing produces a copy of copyrighted material, without explicit approval of the copyright owner, then it is infringing on the copyright.

Where is the unprepared?

Post reply on HN