Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

111–120 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#111
post #103

Earlier quoted context omitted.

We don't seem to be reading the same thing, you're pulling Google out of thin air somewhere.

I'm literally quoting The Verge article and following the links they present...

Must be quoting it wrong then, Google has nothing to do with LLama.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#112
Well theyre claiming their books were scraped illegally from torrents. if i torrent peter pan and watch it alone i can get thrown on jail. If AI is using torrents and getting billions in funding and revenue off a torrented peter pan they should probably be held to the same standard i am.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#113
post #103

Earlier quoted context omitted.

I'm literally quoting The Verge article and following the links they present...

Must be quoting it wrong then, Google has nothing to do with LLama.

Sorry, it's late here. I meant Meta. Thanks for the correction.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#115
post #23
post #21

She'd have to sue every student that writes an essay on a book they'd read

... on a book that they’d illegally acquired then read.

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private tracker.” Bibliotik and the other “shadow libraries” listed, says the lawsuit, are “flagrantly illegal.”

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#116
I was able to overcome the simple "word for word" filtering that is being done on book outputs by prompting ChatGPT to write it in pig latin.

I succeeded getting the first page of Moby Dick - Chapter 1 (Loomings) - Public domain though, but wanted to test.

With ChatGPT primed for pig latin, I also succeeded in getting the first page of Arryhay Otterpay (Book 1) - It happily chattered along ""R.ay andyay Rs.May UrsleyDay, ofay umberNay ourFay, Ivetray riveway, ereway oudpray otay aysay atthay eythay ereway erfectlypay ormalnay, ankthay ouyay eryvay uchmay."

Not perfect pig latin, but that's besides the point.

However, on asking for `Edwetterbay by arahsay ilvermansay`, I faced issues with it citing that is training data didn't include it.

I tried with a book in the same genre ("ieslay hattay helseacay andlerhay oldtay emay"), and ran into the same issue.

When asking about the inconsistency (Why Harry Potter, and not these other books?), it responded: "The excerpt from "Harry Potter and the Philosopher's Stone" that I translated is commonly known and widely referenced, and it's used here as a general example of how a text can be translated into Pig Latin.

For "Lies That Chelsea Handler Told Me", I do not have a widely known or referenced passage from that book in my training data to translate into Pig Latin."

---

TL;DR - I don't think this is cut and dry, but I'm not convinced Silverman has much of a case here.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#117

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.

Agreed, copyright has gone too far. I hope the advent of AI serves to weaken it.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#118
post #77
post #61

Earlier quoted context omitted.

I buy a book and give it to my child, they read the book and later write and sell a story influenced by said book. should that be a copyright infringement? how about they become a therapist and sell access to knowledge from copyrighted books? should that be an infringement? what if they sell access to lectures they've given including facts from said book(s) to millions of people? it's understandable that people feel…

> understand and meet the desires of an audience. LLMs and image generation tech do not do this. For now? I wouldn’t be surprised if that becomes the next feature though.

that's already been a thing for years. long before LLMs and Stable Diffusion

it's just the instagram/tiktok/youtube suggested content algorithms

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#119

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.

Swartz distributed information for everyone to use freely. These companies are processing it privately to develop their for-profit products. Big difference, IMO.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#120

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

> If you downloaded a book from that website, you would be sued and found guilty of infringement.

How often does this actually happen? You might get handed an infringement notice, and your ISP might terminate your service if you're really egregious about it, but I haven't ever heard of someone actually being sued for downloading something.

Post reply on HN