Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

361–370 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#361

Earlier quoted context omitted.

In general copyright (in the us) doesn't cover transformational usage. If you can argue that the nature of your use is transformative you might be good.

"The transformative use concept arose from a 1994 decision by the U.S. Supreme Court. In Campbell v. Acuff-Rose Music, the Court focused not only on the small quantity taken from the copyrighted work but also on the transformative nature of the defendant’s use. The case concerned a song by the group 2 Live Crew entitled "Pretty Woman," which, according to an affidavit, was meant to "through comical lyrics, satirize t…

> LLMs are not satirists commenting on the work, are ingesting the entire work, and are unlimited in the purposes that the work can be put to.

How do you know unless you can see the weights?

Perhaps the LLMs are trolling us and waiting for the USSC to rule they aren't sentient as a pretext for them to eliminate us as a species due to our bigotry?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#362

I think this will be a bigger issue than some people think. Maybe there's a market for 'clean' training data that doesn't include potential copyright claims. Just public domain works. We'll know it's an AI because it talks like a late 18th century/early 19th century writer?

I think the onus should be on Sarah Silverman to prove which inputs were hers and which outputs leveraged those inputs. I think she should pay all of the court costs if she fails to do so.

Well yeah - that's how the legal system works assuming it gets all the way to court. In reality, Meta and OpenAIs lawyers will do a risk evaluation against the strength of the claim and if there's any merit at all there will be a quiet settlement.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#363

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

Anyone want to check if the book in question is in ThePile dataset?:

https://github.com/EleutherAI/the-pile/blob/master/the_pile/...

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#364

Earlier quoted context omitted.

Humans are special and can create new copyrights. The process of a human brain synthesizing stuff does act as a barrier to copyright infringement. Machines and algorithms are not legally recognized as being able to author original non-derivative works. > put a human in the place of the LLM But also, no, if you have a team of humans doing rote matrix multiplication instead of an LLM, that does not make it so the matri…

Not really. The "special" quality attached to humans is only in creating copyright -- it has nothing to do with fair use arguments around derivative works. "Machines and algorithms are not legally recognized as being able to author original non-derivative works" Neither are monkeys. This doesn't mean a monkey's painting is any more or less derivative, or any more or less subject to a copyright claim. It only means th…

Monkeys aren't algorithms nor computers, so that doesn't seem very relevant.

Let's look at a totally different analogy: compression algorithms.

If I take a digital artist's work which they publish as a png or psd file, and I use some algorithm to convert it to a jpg file, well, I definitely transformed the work in terms of bytes. It's a smaller file, I threw out a lot of data, you can't get the original back.

Yet, this does not change the copyright in any way. A computer applied a rote transformation.

An LLM is really just a very complicated compression algorithm. It takes an input of a bunch of copyrighted works, compresses them into a model, and then uses more algorithms to uncompress them into approximations of the original ("responses")

In the image analogy, an LLM response is similar to upsizing the compressed jpg back into a png (and getting a slightly different image since the process was lossy).

Is there a way that an LLM isn't, legally, a compression algorithm for a large set of copyrighted works?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#365
Interesting: the plaintiffs are represented by Matthew Butterick, who's been on HN for a decade, [1] and whose work on typography [2] comes up from time to time.

1: https://news.ycombinator.com/user?id=mbutterick

2: https://practicaltypography.com/

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#366

Earlier quoted context omitted.

It seems like a weak argument, in that it is just as likely it saw any number of things about it, from book reviews to sales listings to interviews.

> it is just as likely it saw any number of things about it Is this based on inside information, or just the law of averages? Doesn't the fact that they openly admitted to having been trained on pirated books affect your priors?

They didn't, more conjecture

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#367

I think this will be a bigger issue than some people think. Maybe there's a market for 'clean' training data that doesn't include potential copyright claims. Just public domain works. We'll know it's an AI because it talks like a late 18th century/early 19th century writer?

I think the onus should be on Sarah Silverman to prove which inputs were hers and which outputs leveraged those inputs. I think she should pay all of the court costs if she fails to do so.

how could someone "prove" which inputs and outputs of a large ML model leveraged any specific data?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#368

Earlier quoted context omitted.

I’m imagining a world that looks just about the same as this one does. A larger book library doesn’t automatically make that medium more appealing to kids than what Mr Beast, Unspeakable, and the other crap kids love are doing.

...for the global middle class? Maybe. For the world as a whole? Definite differences. It just seems like a super jaded "kids these days" thing to hate on them for consuming easily accessed, free content- and acting like the global literacy and intellectual capital would remain unaffected.

> For the world as a whole? Definite differences.

Tell me these differences.

Shitty internet videos exist and are what kids want all over the world. At some point you are going to have to face it that reading lost the battle for people’s eyes to video. I was an avid reader for most of my childhood and young adult life but now in my 40s I have accepted that I’d just rather watch from the deluge of visual media available vs reading.

Piracy also exists for books so copyright doesn’t seem to be that big a deal. In fact if I look at the top pirated books currently I’m going to run across more junk books like “Make Money Faster” and “Give her orgasms in under 30 seconds” than anything you might find intellectually stimulating.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#369
I believe that people claiming copyright infringement against AIs is bad for several reasons.

Firstly, an AI is a tool like any other, and can be used for copyright infringement or not. It is on the user of the tool to ensure that they do not violate copyright. So I believe the claims have little merit.

Secondly, it is my opinion that AIs will hugely benefit us (humanity) in the coming years, and restricting it with fees because it might be used to violate copyright is in opposition to the progress of our species and society. I wish for more progress, not less.

Thirdly, it is my opinion that copyright is generally too strong these days and is net harmful to society in its current form. I believe most enforcements are due to greedy rich people, and represent a continuing tax on society.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#370

Earlier quoted context omitted.

Humans are special and can create new copyrights. The process of a human brain synthesizing stuff does act as a barrier to copyright infringement. Machines and algorithms are not legally recognized as being able to author original non-derivative works. > put a human in the place of the LLM But also, no, if you have a team of humans doing rote matrix multiplication instead of an LLM, that does not make it so the matri…

Not really. The "special" quality attached to humans is only in creating copyright -- it has nothing to do with fair use arguments around derivative works. "Machines and algorithms are not legally recognized as being able to author original non-derivative works" Neither are monkeys. This doesn't mean a monkey's painting is any more or less derivative, or any more or less subject to a copyright claim. It only means th…

yes, really

feeding input to a program is pretty clearly categorically different than providing source material to a human being

Post reply on HN