Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

231–240 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#231
post #164

Earlier quoted context omitted.

Excellent, so you're saying I'll be able to download any copyrighted work from any pirate site and be free of all consequence if I just claim that I'm training an AI?

In general copyright (in the us) doesn't cover transformational usage. If you can argue that the nature of your use is transformative you might be good.

"The transformative use concept arose from a 1994 decision by the U.S. Supreme Court. In Campbell v. Acuff-Rose Music, the Court focused not only on the small quantity taken from the copyrighted work but also on the transformative nature of the defendant’s use. The case concerned a song by the group 2 Live Crew entitled "Pretty Woman," which, according to an affidavit, was meant to "through comical lyrics, satirize the original work." The original work was a rock ballad entitled "Oh, Pretty Woman." The Court was persuaded that no infringement occurred because the defendant added a new meaning and message rather than simply superseding the original work."

So if it is satire, or uses an insignificant piece of the work within a larger work with a different aim or purpose, that's "transformative use," which is something that can be considered when determining "fair use."

LLMs are not satirists commenting on the work, are ingesting the entire work, and are unlimited in the purposes that the work can be put to.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#232

Earlier quoted context omitted.

You can currently download any copyright work from any pirate site and be free of all consequences, and this has always been so. You just can't upload, since that counts as distribution, triggering civil and criminal penalties written in an age before the Internet when only shady commercial operators would distribute unlicensed copyrighted works.

Possession of copyrighted material without permission of the creator is illegal, but yes -- rights holders don't really go after infringers except maybe via ISP three strikes crap. They're very much incentivized to change their behavior for AI scraping, though.

I am not a lawyer, but my understanding is that copyright law typically regulates the unauthorized reproduction, distribution, public display, or creation of derivative works of copyrighted materials. Possession of copyrighted material in itself is not illegal. It's how you use that material that could potentially violate copyright laws.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#233
post #146

Earlier quoted context omitted.

One is a person, the other is a computer program. Legally quite distinct! Note that nobody is even seriously claiming we have an AGI, there's no Star Trek discussion of whether an android is a person. Everyone agrees this is just a computer program.

It doesn't matter if it's a person, or a computer program, or not. This discussion is moot. Is there a substantial reproduction of the works in the output? If not, there's no copyright infringement here. Try reading this legal opinion: https://lawreview.law.ucdavis.edu/issues/53/5/notes/files/53...

Did the trainers illegally access and obtain copyrighted works which are ordinarily protected?

I see multiple questions raised by this suit. Were copyrighted works being illegally stored and distributed by certain sources? Of course they are. Were these illegal sources accessed by the trainers and used to obtain copyrighted works which are not otherwise available for no charge on the public Internet? Are substantial and reproducible copies of these copyrighted works retained within the bowels of the LLM neural nets? Is the LLM able to answer prompts in a way that it would never be able to do, had the copyrighted material not been ingested?

I see the lawsuit addressing several questions at once, and so even a resolution of this suit itself will leave questions unanswered and needing to be kicked upstairs to higher jurisdictions.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#234
post #61

Earlier quoted context omitted.

I'm allowed to make private copies of copywritten works. I'm not allowed to redistribute them. To what extent this is redistribution is not clear. Is there much of difference between this model and a machine, like a VCR, that recreates the original work when I press a button?

I buy a book and give it to my child, they read the book and later write and sell a story influenced by said book. should that be a copyright infringement? how about they become a therapist and sell access to knowledge from copyrighted books? should that be an infringement? what if they sell access to lectures they've given including facts from said book(s) to millions of people? it's understandable that people feel…

> I buy a book and give it to my child

> how about they become a therapist and sell access

> what if they sell access to lectures

I fully agree. But in all of your examples someone is purchasing the right to access the information in question. Did Meta or OpenAI purchase the books (or lectures) with the intention of feeding them into the training for their respective LLM's?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#235
post #181

Earlier quoted context omitted.

>How often does this actually happen? Did you hear about Aaron Schwartz?

He hacked into a server to release a database of paywalled studies to the public. Not only is it not the same but it was the hacking that brought charges upon him.

It's been quite a few years, but AFAIK he didn't hack JSTOR. He downloaded papers en mass using a guest MIT account that had legal access to JSTOR. Maybe he violated their terms, but that is not illegal. He did illegally trespass an unlocked MIT switch closet to do this. They blocked several IPs but his script would rotate to continue. The downloading was over a week or two, enough for security to set up a camera in the closet to catch him retrieving the laptop.

I believe JSTOR sued him to prevent him from releasing the downloaded materials, worried he had offloaded the papers separately from the laptop. The final blow was an outrageous set of charges by the federal government. I also recall several prominent leaders in the open source movement calling it out for what it was, a power trip to make an example of a "digital terrorist". Such a shame.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#236

Earlier quoted context omitted.

Strictly speaking, it's uploading that people get sued for, not downloading. You can download all that you want from Z-Library or BitTorrent, as long as you don't share back. And indexing copyrighted material for search is safe, or at least ambiguous.

Downloading is illegal. That people do not normally get sued or prosecuted for downloading does not mean that they cannot get sued or prosecuted.

It is distribution of copyrighted material without permission of the author that is illegal, when you download you're not distributing so it isn't illegal (unless you're using something like BitTorrent that also distributes it while you're downloading it).

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#237

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

If AI companies get to successfully argue the two points below, what source was used becomes irrelevant. - copyright violation happened before the intervention of the bot - what LLMs spit out is different enough from any of the source that it is not infringing on existing copyright If both stand, I'd compare it to you going to an auction site and studying all the published items as an observer, coming up with your re…

> - copyright violation happened before the intervention of the bot

What is this supposed to mean? The bot didn't "intervene," it was executed by its operators, and it was trained on illicit material obtained by its operators. The LLM isn't on trial. It's not a person.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#238
post #36

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

> Is there any legal basis for saying fair use permits distributing an LLM trained on copyrighted material, but you have to purchase all the content first to do so legally if it's only available for sale? My understanding (disclaimer: IANAL) is that in order to claim fair use, you have to be legally in possession of the work. If the work is only legally available for sale, then you must have legally purchased a copy,…

> in order to claim fair use, you have to be legally in possession of the work.

Which work? The original work, or the derivative work that you're using?

Wikipedia uses non-free content all the time, and they're not purchasing albums to do it. Wikipedia reduces album covers, for example, to low resolution, so that they could not be reused to reproduce a real cover, for example. Sometimes Wikipedia uses screencaps of animated characters, for example, under their non-free content policies. They don't own original copies, they're just hosting low-resolution reproductions. I don't even know what entity would be required to be "legally in possession of the work" for that to be a thing. Could you cite a source, maybe?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#239
post #173

Earlier quoted context omitted.

No, I did. What's your point?

The point is that GP has no reason to believe Google, like Meta, also used copyrighted materials for training its AI. Why did Sarah Silverman sue OpenAI and Meta but not Google?

I didn't accuse Google of using copyrighted materials for training its AI. I accused Google of existing under a different set of laws than mere citizens.

As an example, the mass usage of copyrighted materials to build Youtube.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#240
post #164

Earlier quoted context omitted.

Excellent, so you're saying I'll be able to download any copyrighted work from any pirate site and be free of all consequence if I just claim that I'm training an AI?

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

Can I memorize copyrighted material and recite it on Youtube? What if I do so but imperfectly? Where do you draw the line? If it's infringement for a human to do that why is it not for a LLM?
Post reply on HN