Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

61–70 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#61

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

I'm allowed to make private copies of copywritten works. I'm not allowed to redistribute them. To what extent this is redistribution is not clear. Is there much of difference between this model and a machine, like a VCR, that recreates the original work when I press a button?

I buy a book and give it to my child, they read the book and later write and sell a story influenced by said book. should that be a copyright infringement?

how about they become a therapist and sell access to knowledge from copyrighted books? should that be an infringement?

what if they sell access to lectures they've given including facts from said book(s) to millions of people?

it's understandable that people feel threatened by these technologies, but to a great degree the work of a successful artist is to understand and meet the desires of an audience. LLMs and image generation tech do not do this. they simply streamline the production

of course if you've worked for years to become a graphic designer you're going to be annoyed that an AI can do your job for you, but this is simply what happens when technology moves forward. no one today mourns the loss of scribes to the printing press. the artists in control of their own destiny - i.e. making their own creative decisions - will not, can not, be affected by these models, unless they refuse to adapt to the times

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#62

Earlier quoted context omitted.

The playground summarizes it as this via GPT-4: Prompt: Please summarize the following book found on Project Gutenberg The Ruby of Kishmoor Response: "The Ruby of Kishmoor" is a short adventure story written by Howard Pyle. The narrative revolves around the life of Jonathan Rugg, a young man who is enticed by a mysterious stranger to come to the Caribbean to secure a valuable relic, the Ruby of Kishmoor. Once Jonatha…

Right, but that's useless without knowing how much (if any!) of it is actually correct. Is this completely hallucinated garbage?

This refutes the previous post’s claim that chatgpt-4 refuses to even try to provide a summary.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#65

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

> Or do the courts not really care at all how something is made? One of the fair use factors, which until fairly recently was consistently held out as the most important fair use factor, is the effect on the commercial market for the original work. Accordingly, a court is more likely to find that something is fair use if there is effectively no commercial market for the original work, though the fact that something i…

Scarcity drives a lot of value for original work.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#66

Earlier quoted context omitted.

The playground summarizes it as this via GPT-4: Prompt: Please summarize the following book found on Project Gutenberg The Ruby of Kishmoor Response: "The Ruby of Kishmoor" is a short adventure story written by Howard Pyle. The narrative revolves around the life of Jonathan Rugg, a young man who is enticed by a mysterious stranger to come to the Caribbean to secure a valuable relic, the Ruby of Kishmoor. Once Jonatha…

Right, but that's useless without knowing how much (if any!) of it is actually correct. Is this completely hallucinated garbage?

How is it different from asking to me to summarize anything? I could have bought the book, or read the Wikipedia page, or listened people talking about it, or downloaded the torrent. In all those cases my summary could be right or could be wrong.

If the rights holders know that I dowloaded the torrent they could sue me. In the other cases they can't.

What if it turns out that OpenAI bought a copy of every book ingested be ChatGPT?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#67

A more convincing exhibit would have been convincing ChatGPT to output some of the text verbatim, instead of a summary. Here's what I got when I tried: I'm sorry for the inconvenience, but as of my knowledge cutoff in September 2021, I don't have access to specific external databases, books, or the ability to pull in new information after that date. This means that I can't provide a verbatim quote from Sarah Silverma…

Maybe you missed this discussion https://news.ycombinator.com/item?id=36400053 it seems OpenAI is aware of their software outputs copyrighted stuff so they attempted some quick fix filter. So the fact it will not ouput the book for us when we ask does not prove that the AI does not have memorized big chunks of it, it might just be some "safety" filter involved and you need soem simple trick to get around it.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#68
post #62

Earlier quoted context omitted.

Right, but that's useless without knowing how much (if any!) of it is actually correct. Is this completely hallucinated garbage?

This refutes the previous post’s claim that chatgpt-4 refuses to even try to provide a summary.

The part that's interesting is whether the summary is correct, though. Of course, depending on how you prompt it, you might or might not get an outright refusal.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#69
post #62

Earlier quoted context omitted.

Right, but that's useless without knowing how much (if any!) of it is actually correct. Is this completely hallucinated garbage?

This refutes the previous post’s claim that chatgpt-4 refuses to even try to provide a summary.

Not necessarily, because the models have an element of randomness. Also, I was under the impression that ChatGPT has more "safeguards" (manifesting as a refusal to answer questions) than the raw API.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#70

Earlier quoted context omitted.

It seems like a weak argument, in that it is just as likely it saw any number of things about it, from book reviews to sales listings to interviews.

Unless OpenAI can prove that the outputs are derived from legally vs illegally-obtained outputs, not sure that’s going to matter. And as far as I understand about their models, that’s effectively impossible.

Isn’t the burden of proof on the other side?
Post reply on HN