Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

211–220 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#211

Earlier quoted context omitted.

> I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. I would be very happy if either a court or lawmakers decided that copyright itself was unconscionable. That isn't what's going to happen, though. And I think it's incredibly unacceptable if a court…

It is the equivalent of making a 3D map of a museum and getting sued by one artist of one painting in the museum. Ant individual work in an AI dataset is nearly worthless - only in aggregate does it have value. If that doesn't count as a "transformative work" I don't know what does.

> It is the equivalent of making a 3D map of a museum and getting sued by one artist of one painting in the museum.

If the painting is copyrighted (rather than public domain, as many pieces in museums are), and the map includes an image of that painting, I would expect that to be prohibited. I would prefer the world in which copyright doesn't exist, but while it exists, it should apply to everyone equally.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#212

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

If AI companies get to successfully argue the two points below, what source was used becomes irrelevant.

- copyright violation happened before the intervention of the bot

- what LLMs spit out is different enough from any of the source that it is not infringing on existing copyright

If both stand, I'd compare it to you going to an auction site and studying all the published items as an observer, coming up with your research result, to then be sued because some of the items were stolen. Going after the theaves make sense, does going after the entity that just looked at the stolen goods make sense ?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#213
post #66

Earlier quoted context omitted.

Right, but that's useless without knowing how much (if any!) of it is actually correct. Is this completely hallucinated garbage?

How is it different from asking to me to summarize anything? I could have bought the book, or read the Wikipedia page, or listened people talking about it, or downloaded the torrent. In all those cases my summary could be right or could be wrong. If the rights holders know that I dowloaded the torrent they could sue me. In the other cases they can't. What if it turns out that OpenAI bought a copy of every book ingest…

> What if it turns out that OpenAI bought a copy of every book ingested be ChatGPT?

That still doesn't necessarily confer to them the right to use it to train a model and generate derivative works based on purchased content.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#214

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Strictly speaking, it's uploading that people get sued for, not downloading. You can download all that you want from Z-Library or BitTorrent, as long as you don't share back. And indexing copyrighted material for search is safe, or at least ambiguous.

Downloading is illegal. That people do not normally get sued or prosecuted for downloading does not mean that they cannot get sued or prosecuted.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#215
post #83
post #38

>On information and belief, the reason ChatGPT can accurately summarize a certain copyrighted book is because that book was copied by OpenAI and ingested by the underlying OpenAI Language Model (either GPT-3.5 or GPT-4) as part of its training data. While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false. First, even if the mode…

> this quote from the complaint seems obviously false I notice you go on to provide an argument only for why it might not be true. Also, seeing the other post on this, I asked chatgpt-4 for a summary of “ The Ruby of Kishmoor” as well, and it provided one to me, though I had to ask twice. I don’t know anything about that book, so I can’t tell if its summary is accurate, but so much for your test. It seems pretty naiv…

> IMO, a better argument is that this is fair use

There is no way in Hell that this is fair use!

Fair use defenses rest on the fact that a limited excerpt was used for limited distribution, among other criteria.

For example, if I'm a teacher and I make 30 copies of one page of a 300-page novel and I hand that out to my students, that's a brief excerpt for a fairly limited distribution.

Now if I'm a social media influencer and I copy all 300 pages of a 300-page book and then I send it out to all 3,000 of my followers, that's not fair use!

Also if I'm a teacher, and I find a one-page infographic and I make 30 copies of that, that's not fair use, because I didn't make an excerpt but I've copied 100% of the original work. That's infringement now.

So if LLMs went through en masse in thousands of copyrighted works in their entirety and ingested every byte of them, no copyright judge on the planet would call that fair use.

For reference, the English Wikipedia has a policy that allows some fair-use content of copyrighted works: https://en.wikipedia.org/wiki/Wikipedia:Non-free_content_cri...

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#216
post #139

Earlier quoted context omitted.

Because LLMs are not people. They are nothing like people; not in construction nor behaviour.

LLMs are built upon neural networks which are modelled upon how brains work can you explain to me specifically how they're different? can you explain to me how they're different to the degree that making an analogy between the two is "disingenuous to the extreme"?

> LLMs are built upon neural networks which are modelled upon how brains work

You are confused. Neural networks are inspired by how brains work, but they do not actually simulate brains.

Airplanes are also inspired by how birds work, but (presumably) you don't think that bird laws should apply to airplanes.

> can you explain to me how they're different to the degree that making an analogy between the two is "disingenuous to the extreme"?

It's disingenuous because you don't believe that either.

If you think that LLMs are just like human brains, and should be allowed to learn from books the same as people, then presumably you also believe they are entitled to all other human rights: to vote, to live, &c. If you operate an LLM and you shut it down then that's murder and you belong in jail.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#217
post #183

Earlier quoted context omitted.

analogies are not descriptions of the things themselves, otherwise they would not be analogies, would they? now remember that this is an analogy . re-read my comments in this light and perhaps we can continue this conversation in a more grounded and reasonable manner however, I'll be frank: have you studied neural networks? if you haven't, it's very difficult to take you seriously on this

Now you’re just being condescending. I did read your comments and I know what an analogy is. Consider that there is a different perspective to yours that can validly view your analogy as absurd. Yes I have studied neural nets and have a good understanding of their function. I am still not sure how, despite their development being inspired by animal brains, you can liken an LLM to an actual person. There are so many v…

your argument was predicated upon my comment being factually inaccurate, not analogously poor. I merely listed some leading questions analogising the ability of human brains to absorb, contextualise and emit copyrighted information to an LLM's ability to do all of those same things. you created a weaker position to attack which was that LLMs are the same as humans

if you want to discuss whether it is a poor analogy or not, I'm all for that, but the passage you've chosen to go down is to act as if it were not an analogy at all, which—to borrow a phrase—is disingenuous to the extreme

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#218
post #50

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

Eventually, I imagine a new licensing concept will emerge, similar to the idea of music synchronization rights -- maybe call it "training rights." It won't matter whether the text was purchased or pirated -- just like it doesn't matter now if an audio track was purchased or pirated, when it's mixed into in a movie soundtrack. Talent agencies will negotiate training rights fees in bulk for popular content creators, wh…

I suspect the opposite outcome also being plausible: the LLM is viewed analogously to a blog author. The blogger/LLM may consume a book, subsequently produce "derived" output (generated text), and thus generate revenue for the blogger/LLM's employer. Consequently, the blogger/LLM's output -- while "derived" in some sense -- differs enough to be considered original work, rather than "derivative work" (like a book's film adaptation). Auditing how the blogger/LLM consumed relevant material is thus absurd.

Of course, this line of reasoning hinges on the legitimacy of an "LLM agent blogger agent" type of analogy. I suspect the equivalence will become more natural as these AI agents continue to rapidly gain human-like qualities. How acceptable that perspective would be now, I have no idea.

In contrast, if the output of a blogger is legally distinct from an AI's, the consequences quickly become painful.

* A contract agency hires Anne to practice play recitals verbally with a client. Does the agency/Anne owe royalties for the material they choose? What if the agency was duped, and Anne used -- or was -- a private AI which did everything?

* How does a court determine if a black box AI contains royalty-requiring training material? Even if the primary sources of an AI's training were recorded and kosher, a sufficiently large collection of small quotes could be reconstructed into an author's story.

* What about AIs which inherit (weights, or training data generated) from other AIs of unknown training provenance? Or which were earlier trained on some materials with licenses that later changed? Or AIs that recursively trained their successors using copyrighted works which it AI reconstructed from legal sources? When do AIs become infected with illegal data?

The business of regulating learning differently depending on whether the agent uses neurons or transistors seems...fraught. Perhaps there's a robust solution for policing knowledge w.r.t silicon agents. If you have an idea, please share!

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#219
post #164

Earlier quoted context omitted.

Machine learning models have been trained with copyrighted data for a long time. Imagenet is full of copyrighted images, clearview literally just scanned the internet for faces, and I am sure there are other, older examples. I am unsure if this has been tested as fair use by a US court, but I am guessing it will be considered to be so if it is not already.

Excellent, so you're saying I'll be able to download any copyrighted work from any pirate site and be free of all consequence if I just claim that I'm training an AI?

In general copyright (in the us) doesn't cover transformational usage. If you can argue that the nature of your use is transformative you might be good.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#220
post #146

Earlier quoted context omitted.

One is a person, the other is a computer program. Legally quite distinct! Note that nobody is even seriously claiming we have an AGI, there's no Star Trek discussion of whether an android is a person. Everyone agrees this is just a computer program.

It doesn't matter if it's a person, or a computer program, or not. This discussion is moot. Is there a substantial reproduction of the works in the output? If not, there's no copyright infringement here. Try reading this legal opinion: https://lawreview.law.ucdavis.edu/issues/53/5/notes/files/53...

I don't believe it's copyright infringement either, but I don't believe it should be allowed either.

It used to be that when people bought a piece of land they owned that land from the center of the Earth to the end of the Universe. When planes were invented the laws on who owned the skies had to change or planes wouldn't work.

It's the same here, but in reverse. If LLMs aren't prohibited from learning without permission then people will be forced to hide their works.

Post reply on HN