Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

71–80 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#71
post #62

Earlier quoted context omitted.

Right, but that's useless without knowing how much (if any!) of it is actually correct. Is this completely hallucinated garbage?

This refutes the previous post’s claim that chatgpt-4 refuses to even try to provide a summary.

chatgpt-4 != gpt4 on the openAI playground

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#72
post #66

Earlier quoted context omitted.

Right, but that's useless without knowing how much (if any!) of it is actually correct. Is this completely hallucinated garbage?

How is it different from asking to me to summarize anything? I could have bought the book, or read the Wikipedia page, or listened people talking about it, or downloaded the torrent. In all those cases my summary could be right or could be wrong. If the rights holders know that I dowloaded the torrent they could sue me. In the other cases they can't. What if it turns out that OpenAI bought a copy of every book ingest…

> In all those cases my summary could be right or could be wrong.

Well that's incredibly nihilistic. Whether the summary is correct or not matters a great deal! And if someone I knew said they read a book, even a very obscure one, and then summarized it to me, I'd have great confidence that they would get such simple facts as "who are the characters" and "what are the major plot points" correct.

But ChatGPT? Who the hell knows? You can't trust a thing it says, especially about obscure topics. The summary is useless if you have to do a bunch of verification to see if any of it is even true, a problem that summaries even by moderately competent human writers don't have!

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#73
Sarah's pov raises some questions for me regarding my own "training", there is a noteworthy part of who I am built upon the consumed music, books, movies, video games and tv shows that myself or people around me have pirated and shared with me. This part of me helped me in life appreciably, I could also say I profited because of it, helping me along my life in being likable, funny, relatable, with broad outlooks etc.

Is my brain just by the act of existing continually infringes on copyright? Can I be sued because I made a reference to a movie I pirated or because I whistle a song I never bought?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#74

Earlier quoted context omitted.

Unless OpenAI can prove that the outputs are derived from legally vs illegally-obtained outputs, not sure that’s going to matter. And as far as I understand about their models, that’s effectively impossible.

Isn’t the burden of proof on the other side?

Not when OpenAI publicly declared they trained on pirated works. I can’t imagine “we can’t tell if this is the result of the illegal thing we did or not” is going to stand up very well, nor does it bode well for any refutation of the plaintiff’s depiction of their intent. Part of fair use consideration is commercial impact and when you steal a bunch of books to train your AI model, it’s hard to refute that the impact is not negative or that you didn’t intend commercial harm.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#75
>> In the OpenAI suit, the trio offers exhibits showing that when prompted, ChatGPT will summarize their books, infringing on their copyrights.

This doesn't seem like copyright infringement. I could read the book and offer a summary right? Someone on goodreads could as well. Why should an AI doing it be different? BTW I could also read someone's illicit copy and do the same, couldn't I?

I think people are trying to claim exclusive use rights that they simply don't have. I look forward to a lawyers opinion on this one.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#76
post #10

I severely doubt that the spiders which crawled the data would go to the trouble of dereferencing and downloading torrents.

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private tracker.” Bibliotik and the other “shadow libraries” listed, says the lawsuit, are “flagrantly illegal.”

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#77
post #61

Earlier quoted context omitted.

I'm allowed to make private copies of copywritten works. I'm not allowed to redistribute them. To what extent this is redistribution is not clear. Is there much of difference between this model and a machine, like a VCR, that recreates the original work when I press a button?

I buy a book and give it to my child, they read the book and later write and sell a story influenced by said book. should that be a copyright infringement? how about they become a therapist and sell access to knowledge from copyrighted books? should that be an infringement? what if they sell access to lectures they've given including facts from said book(s) to millions of people? it's understandable that people feel…

> understand and meet the desires of an audience. LLMs and image generation tech do not do this.

For now? I wouldn’t be surprised if that becomes the next feature though.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#78
post #61

Earlier quoted context omitted.

I'm allowed to make private copies of copywritten works. I'm not allowed to redistribute them. To what extent this is redistribution is not clear. Is there much of difference between this model and a machine, like a VCR, that recreates the original work when I press a button?

I buy a book and give it to my child, they read the book and later write and sell a story influenced by said book. should that be a copyright infringement? how about they become a therapist and sell access to knowledge from copyrighted books? should that be an infringement? what if they sell access to lectures they've given including facts from said book(s) to millions of people? it's understandable that people feel…

If you found a way to have a million children who could grow up in one day your analogy would be more apt. In that case you and your children would rightly be considered a threat.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#79
post #66

Earlier quoted context omitted.

Right, but that's useless without knowing how much (if any!) of it is actually correct. Is this completely hallucinated garbage?

How is it different from asking to me to summarize anything? I could have bought the book, or read the Wikipedia page, or listened people talking about it, or downloaded the torrent. In all those cases my summary could be right or could be wrong. If the rights holders know that I dowloaded the torrent they could sue me. In the other cases they can't. What if it turns out that OpenAI bought a copy of every book ingest…

I feel like thats one of the many questions regulators and law makers are going to be asked long term. I'm sure buying the book for "commercial purposes" like that would't be appropriate, but then again, does that mean if I read it and then summarize it in my work, or regurgitate its info as part of my job...I'm violating a license?

A world where humans have special permissions but LLMs don't seems pretty interesting to consider, especially if they're both doing the same kind of things with the data.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#80

>> In the OpenAI suit, the trio offers exhibits showing that when prompted, ChatGPT will summarize their books, infringing on their copyrights. This doesn't seem like copyright infringement. I could read the book and offer a summary right? Someone on goodreads could as well. Why should an AI doing it be different? BTW I could also read someone's illicit copy and do the same, couldn't I? I think people are trying to c…

I think the argument is somewhat more interesting if the book was pirated--both you as an individual and OpenAI as a company could be sued for that.

But I really don't see how you could prove OpenAI did that, since ChatGPT could have learned from existing summaries on Wikipedia and Goodreads.

Post reply on HN