Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

321–330 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#321
post #79

Earlier quoted context omitted.

I feel like thats one of the many questions regulators and law makers are going to be asked long term. I'm sure buying the book for "commercial purposes" like that would't be appropriate, but then again, does that mean if I read it and then summarize it in my work, or regurgitate its info as part of my job...I'm violating a license? A world where humans have special permissions but LLMs don't seems pretty interesting…

There's nothing illegal about reading a book and then circulating your summary/review of it. This doesn't even get into issues of fair use because you aren't redistributing any of the copyrighted material in the first place, merely facts and your own opinions about it. The separate issue that's concerning is that GPT can't be trusted to accurately summarize anything obscure at all, but it'll sure throw text at you no…

> There’s nothing illegal about reading a book and then circulating your summary/review of it. This doesn’t even get into issues of fair use because you aren’t redistributing any of the copyrighted material in the first place, merely facts and your own opinions about it.

A summary may or may not be a derivative work before considering Fair Use; “redistributing copyrighted material” isn’t the only exclusive right of copyright: producing copies is, but more to the point so is producing derivative works.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#322
post #97

Earlier quoted context omitted.

I don’t doubt the poster was telling the truth when they said they asked for a summary of the book and didn’t get one. It refutes the idea that chatgpt’s inability to provide a summary means it didn’t scan the original text: since it can provide a summary, the argument is entirely spurious.

It's also a silly test since there is undoubtably a summary of most books someplace on the web.

You'd be surprised. There are an ungodly number of books published every year, and the back catalog is huge. You're over-indexing on popular books.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#323

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

> copyright is the reason ~every child doesn't have access to ~every book ever written.

And? Is there some reason anybody, child or adult, deserves access to "every" anything? Should children have access to every video game ever made, every Matchbox car, every Lego set?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#324

Earlier quoted context omitted.

Or maybe we could figure out a new economic model, instead of blindly sticking with one based on the limitations of the pre-digital age.

How about you go figure out this new economic model, and come back when it's ready. Until then, the existing model will persist, thank you

First change the incentives, then everyone will work on finding a new economic model.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#326
post #282

Earlier quoted context omitted.

This isn't completely new, similar issues came up with search engines and this may be seen as 'transformative'. But there may be issues with models that happily reproduce copyrighted texts in their entirety along with other novel issues like models that hallucinate defamatory things or other such problems. Still, I doubt this particular genie can be stuffed back into the bottle, so we'll probably see a lot of litigat…

I agree it's not an entirely new issue. But it's a little different from search results. Say I use the generative paint brush in photoshop. It reproduces a portion of the copyrighted work. I then use the image on an advertising campaign, other merchandise, or post the final product as my own work. Would I be responsible? Would Adobe? Given that retraining these models is not simple, or cheap, would this be just 'cost…

This already sort of happened in 2008 when Chuck Close forced someone to stop distributing a Photoshop plug-in that imitated his style.

https://hyperallergic.com/54104/my-chuck-close-problem/

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#327

> The lawsuit against OpenAI alleges that summaries of the plaintiffs’ work generated by ChatGPT indicate the bot was trained on their copyrighted content. “The summaries get some details wrong” but still show that ChatGPT “retains knowledge of particular works in the training dataset," the lawsuit says. Setting aside the whole issue of whether LLM constitutes a derived work of whatever it's trained on, this sounds l…

That isn't firm evidence, but courts don't need firm evidence to start a case and discover new facts. They very well can ask LLM experts, and openAI themselves, whether that output is highly likely to have been derived from the copyrighted work in question. Anyway. If the argument is "No, it's not from the book, it's from someone else's copyrighted summary", that just means the person who wrote such a summary needs t…

A summary can be written in such a way as to violate copyright itself. So even if they say "We trained it on the following summaries:...," there could be an issue.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#328

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

If AI companies get to successfully argue the two points below, what source was used becomes irrelevant. - copyright violation happened before the intervention of the bot - what LLMs spit out is different enough from any of the source that it is not infringing on existing copyright If both stand, I'd compare it to you going to an auction site and studying all the published items as an observer, coming up with your re…

I'd argue that if an automated process can ingest A and spit out B, then B is inherently a derivative work of A. (Never mind that humans are also automata.)

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#329
post #282

Earlier quoted context omitted.

This isn't completely new, similar issues came up with search engines and this may be seen as 'transformative'. But there may be issues with models that happily reproduce copyrighted texts in their entirety along with other novel issues like models that hallucinate defamatory things or other such problems. Still, I doubt this particular genie can be stuffed back into the bottle, so we'll probably see a lot of litigat…

I agree it's not an entirely new issue. But it's a little different from search results. Say I use the generative paint brush in photoshop. It reproduces a portion of the copyrighted work. I then use the image on an advertising campaign, other merchandise, or post the final product as my own work. Would I be responsible? Would Adobe? Given that retraining these models is not simple, or cheap, would this be just 'cost…

In that particular case if you have an enterprise licence, Adobe have accepted responsibility:

>Adobe is so confident its Firefly generative AI won’t breach copyright that it’ll cover your legal bills The offer is available only to users of its enterprise Firefly product, which launches today

https://www.fastcompany.com/90906560/adobe-feels-so-confiden...

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#330

Earlier quoted context omitted.

How about you go figure out this new economic model, and come back when it's ready. Until then, the existing model will persist, thank you

First change the incentives, then everyone will work on finding a new economic model.

This makes even less sense than the previous guy. When you've figured out how the world will work without people incentivized by money and power, get back to the rest of us
Post reply on HN