Live data from Hacker News

Penguin Random House underscores copyright protection in AI rebuff

thebookseller.com

41–50 of 57 posts

Re: Penguin Random House underscores copyright protection in AI rebuff

#41

Earlier quoted context omitted.

I think that’s a more interesting question. I’m not sure how to do it, but I think that finding a way to stop the reproduction of copyrighted content is probably the missing piece. If there was a monetary penalty for reproduction of copyright works (apply human laws to machine), then I bet these companies would quickly figure out how to fingerprint output and match it to source data before sending it to the user.

So I go into a cinema with my recording equipment, record the movie, step outside and start selling downloads, and when the cops come, I say: "I’m not sure how to do it, but I think that finding a way to stop the reproduction of copyrighted content is probably the missing piece." Really?

That's not a close reading of the other comment.

If LLMs were specially advertised as a way to get all the stuff you already love for free, it would be.

They're not.

What they are advertised as is a way to solve problems and create novel things.

In this regard, LLMs are less of a copyright infringement issue than, e.g. Google News and Google Images both of which had to change because they were in law copyright infringements.

This does not mean that LLMs or diffusion models must get a free pass or anything like that: these are a novel things that didn't previously exist, so while I was initially surprised by the legal cases against Stability and OpenAI due to the existence of Google search and that nobody seemed to care about GPT-3 or the original DALL•E, the arguments made against them are nevertheless interesting and worth caring about.

I suspect LLMs are so useful the powers that be will just carve out a space for them; conversely I don't see that applying to image generation models, so I kinda expect them to be strictly limited to cases where the source data can be proven correctly licensed.

Re: Penguin Random House underscores copyright protection in AI rebuff

#42
post #24

Earlier quoted context omitted.

> The copyright violation is in its use for training data. Every textbook, every educational TV show or YouTube video, every artist whose works have been shown to me in my school lessons in art, music, graphic design, usw. have all been copyrighted. With one exception: Shakespeare. Even the bible, being a translation*, had copyright notices on it. But worse than that: if it were so even for scraping the whole web and…

I agree that copyright law is woefully unprepared, and that LLMs have become useful enough it will be difficult if not impossible to stop their development. > Page Rank is a big ol' matrix multiplication, and it spits out quotes verbatim. That makes no sense. Google Books was blocked exactly because it was reproducing enough material to be infringing. Google search usually does not copy enough material that it might…

> That makes no sense. Google Books was blocked exactly because it was reproducing enough material to be infringing. Google search usually does not copy enough material that it might be considered infringing (even though I think news sites have sued over this in some jurisdictions, right?)

Google Books? That's a non-sequitur. When I said "Page Rank" I mean the algorithm used by Google Search to rank web pages in their search engine results.

The search results give you direct quotations from the pages they linked to, and for a long time also had links to cached copies of those pages.

> It's fairly obvious that the people running LLMs don't want to prevent them from infringing copyright, particularly in image gen.

Apart from all the times they pop up messages refusing to reproduce copyrighted material? The only reason I was even willing to create the following query is because I knew it would refuse me on that grounds: https://chatgpt.com/share/67138f5b-a110-8011-adf8-a82f3fc473...

Image gen… well, for at least the third-party fine-tuners, I agree with your impression. The main players? Unclear to me, especially as we see some models from big players which are created from scratch using only correctly licensed materials — I believe Adobe's models would be an example of correctly licensed training content.

Re: Penguin Random House underscores copyright protection in AI rebuff

#43

Earlier quoted context omitted.

I disagree: I think copyright law is wholly prepared. If some thing produces a copy of copyrighted material, without explicit approval of the copyright owner, then it is infringing on the copyright. Where is the unprepared?

It's unprepared for the fact that you can now build thinking machines using these new techniques, and that governments will be reluctant to regulate them if they think it will give them issues in an AI race against strategic opponents. Of course it's doubtful that image and video gen AI are strategically important, but I trust the AI lobby to make sure that governments won't make that distinction.

> Of course it's doubtful that image and video gen AI are strategically important, but I trust the AI lobby to make sure that governments won't make that distinction.

I suspect otherwise. The use cases for image generators are much less obvious than for a personal assistant that's just as fluent in Chinese, international relations, hacking, military strategy, and speech writing… even when the current fluency of LLMs relative to government workers is somewhere between "mediocre" to "it'll do".

I think big copyright owners, however, will probably make their own models. Disney etc. could very easily use their own content to (help) make as many more films as they want. Real human artists will be disempowered regardless of how the copyright stuff happens. Probably.

Re: Penguin Random House underscores copyright protection in AI rebuff

#44
post #43

Earlier quoted context omitted.

It's unprepared for the fact that you can now build thinking machines using these new techniques, and that governments will be reluctant to regulate them if they think it will give them issues in an AI race against strategic opponents. Of course it's doubtful that image and video gen AI are strategically important, but I trust the AI lobby to make sure that governments won't make that distinction.

> Of course it's doubtful that image and video gen AI are strategically important, but I trust the AI lobby to make sure that governments won't make that distinction. I suspect otherwise. The use cases for image generators are much less obvious than for a personal assistant that's just as fluent in Chinese, international relations, hacking, military strategy, and speech writing… even when the current fluency of LLMs…

I hope so, but big AI is banking on being able to use those models for everything.* Our "luck" is that Disney and other giant corporate copyright holders will indeed sue them to oblivion, or at least try.

* Afaik image gen AI is one the few profitable use cases, with Midjourney apparently being profitable already.

Re: Penguin Random House underscores copyright protection in AI rebuff

#45
post #19

Penguin Random House is a predatory corrupt business that the world would be much better off without. If AI means businesses like this shut down I need more „AI“

Honest question, what's "predatory corrupt" about their business? I have some of their paperbacks on my bookshelves and they've been a pretty decent value for money thing.

All major publishers are known to be fundamentally rent-seeking entities. Their vast history of litigative bullying and anti-consumer behavior has created a self-serving perception of them as only capable of being high-profile scoundrels. Which means even if there somehow exist some good ones in their current iteration, they will always be evil for many people just by being associated with that domain.

Re: Penguin Random House underscores copyright protection in AI rebuff

#47

Which country has the most favorable 'fair use' laws, and why wouldn't big companies train their models there?

Would that matter if the company wants to do business in countries with more restrictive laws? I.E. if I wrote my own spin-off of a popular book series, which was somehow considered fair use in country A, but considered infringing in country B, the publisher could get it removed from stores in country B. By the same logic, if AI training is ruled as copyright infringement in the US, it won't matter if the company tra…

IANAL but the article has a quotation from a lawyer that says that the infringing act is the training.

Re: Penguin Random House underscores copyright protection in AI rebuff

#48
post #43

Earlier quoted context omitted.

> Of course it's doubtful that image and video gen AI are strategically important, but I trust the AI lobby to make sure that governments won't make that distinction. I suspect otherwise. The use cases for image generators are much less obvious than for a personal assistant that's just as fluent in Chinese, international relations, hacking, military strategy, and speech writing… even when the current fluency of LLMs…

I hope so, but big AI is banking on being able to use those models for everything.* Our "luck" is that Disney and other giant corporate copyright holders will indeed sue them to oblivion, or at least try. * Afaik image gen AI is one the few profitable use cases, with Midjourney apparently being profitable already.

Re image gen cost, by my estimate inference was $0.0001/image 2 years ago: https://benwheatley.github.io/blog/2022/10/09-19.33.04.html

I suspect that there are significantly better options than the thing which just happened to be accessible to me at the time, so I would say that quality, not cost, is the thing currently preventing feature length movies from being generated for less than the cost a movie theatre ticket.

Re: Penguin Random House underscores copyright protection in AI rebuff

#49

Earlier quoted context omitted.

> The AI companies contend that it's fair use Do they? "Fair Use" is an affirmative defense, so the only time we're going to get into that is in a court case, where it'll be tested through legal means. I would say it's even more nuanced: if LLM training involves merely reading a dataset, but it is not strictly necessary to copy , or even store it verbatim to be useful, then does it even fall under copyright protectio…

> if LLM training involves merely reading a dataset, but it is not strictly necessary to copy, or even store it verbatim to be useful, then does it even fall under copyright protection at all? Copyright includes the creation of derivative works, not just literally copying the source material. For instance, imagine I read a novel, then I decide to write my own, unauthorized sequel to it. It's not a literal "copy" of t…

LLM-generated book clones (as seen on Amazon and elsewhere) could potentially fall afoul of copyright law in many ways, including: rights to reproduction/substantial similarity; derivative works; adaptation (including translation); distribution; performance and public display (including broadcast or transmission); etc.

Re: Penguin Random House underscores copyright protection in AI rebuff

#50

Earlier quoted context omitted.

> if LLM training involves merely reading a dataset, but it is not strictly necessary to copy, or even store it verbatim to be useful, then does it even fall under copyright protection at all? Copyright includes the creation of derivative works, not just literally copying the source material. For instance, imagine I read a novel, then I decide to write my own, unauthorized sequel to it. It's not a literal "copy" of t…

LLM-generated book clones (as seen on Amazon and elsewhere) could potentially fall afoul of copyright law in many ways, including: rights to reproduction/substantial similarity; derivative works; adaptation (including translation); distribution; performance and public display (including broadcast or transmission); etc.

LLMs don't necessarily need to reproduce their source material to make use of it. They could summarize, analyze, condense, paraphrase, extract statistics or factoids. There's also the question of how the models actually store the source material or not. It's physically impossible for the verbatim text to live in the model weights, and so at the very least, it's compressed or abstracted. So any copyright claims will need to get beyond a simplistic allegation of copying, for sure.
Post reply on HN