Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

421–430 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#421

Earlier quoted context omitted.

Or maybe we could figure out a new economic model, instead of blindly sticking with one based on the limitations of the pre-digital age.

How about you go figure out this new economic model, and come back when it's ready. Until then, the existing model will persist, thank you

Oh boy, do I have good news for you!

There are already many writers making thousands of dollars a month by publishing free serialized web novels, via Patreon. Some are using their own websites, but most are on Royal Road (or scribblehub, webnovel, wattpad, AO3).

A random example from Royal Road[1], the author makes $12065/month. Mind you, the text is not gated, it's free to read, the patreon only offers early access...

[1] https://www.royalroad.com/fiction/63759/super-supportive [2] https://www.patreon.com/Sleyca

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#422
post #292

Earlier quoted context omitted.

IANAL either, but FWIW, it's literally in the name - copy right. Not "ownership rights", but "copying rights".

That's not terribly relevant for Internet applications, because for the most part people deliberately cause the computer they control to download and save copyrighted material to storage they control, and then consume it at leisure. That's copying.

If I understand right (I'm not a lawyer), something like this came up in a case about a utility for cheating in an online game, and another much older one about the copy of an app made in memory in order to run it.

For the purpose for which the software was sold and bought, the in-memory copy is legit[1]. For cheating, it's a copyright violation[0].

[0] https://www.engadget.com/2008-07-15-blizzard-wins-lawsuit-ag...

[1] http://digital-law-online.info/lpdi1.0/treatise20.html

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#423
post #61

Earlier quoted context omitted.

I buy a book and give it to my child, they read the book and later write and sell a story influenced by said book. should that be a copyright infringement? how about they become a therapist and sell access to knowledge from copyrighted books? should that be an infringement? what if they sell access to lectures they've given including facts from said book(s) to millions of people? it's understandable that people feel…

> I buy a book and give it to my child > how about they become a therapist and sell access > what if they sell access to lectures I fully agree. But in all of your examples someone is purchasing the right to access the information in question. Did Meta or OpenAI purchase the books (or lectures) with the intention of feeding them into the training for their respective LLM's?

this I do not know. it's not the aspect I'm exploring really

however, in the analogy the books could have been read for free from a library

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#424

Earlier quoted context omitted.

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

Using charged language of bringing children into the equation is not a good way in having a discussion.

...On the other hand, seems to work gangbusters for getting policy passed.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#425

A more convincing exhibit would have been convincing ChatGPT to output some of the text verbatim, instead of a summary. Here's what I got when I tried: I'm sorry for the inconvenience, but as of my knowledge cutoff in September 2021, I don't have access to specific external databases, books, or the ability to pull in new information after that date. This means that I can't provide a verbatim quote from Sarah Silverma…

GPT is a lossy jpeg of the whole Internet. It’s not possible to extract verbatim text from it, due to how neural networks work.

How do you think they would fit exabytes of text data into a gigabyte-sized neural network? That’s right, it’s lossy.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#426

Generating new text inspired by and iterated from concepts of older works is pretty much how human beings write. I understand that nobody wants a multi billion dollar company to take their works without giving compensation, but I worry all of this could move dangerously close to allowing concepts and styles to be copyrighted and crack down on transformative works.

The problem is not that they summarize text, the problem is the regular old copyright infringement that happens before training the model.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#427
post #193

Well theyre claiming their books were scraped illegally from torrents. if i torrent peter pan and watch it alone i can get thrown on jail. If AI is using torrents and getting billions in funding and revenue off a torrented peter pan they should probably be held to the same standard i am.

I don't think you can be thrown in jail. Just torrenting a film and watching it would be a civil offense? Now making copies and selling them on the street corner is another story.

Depending on your jurisdiction, it may be a civil offense. Where I’m at, it’s still a criminal offense albeit one that would probably never be prosecuted because the fine for one infringement would be minuscule.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#428

I was able to overcome the simple "word for word" filtering that is being done on book outputs by prompting ChatGPT to write it in pig latin. I succeeded getting the first page of Moby Dick - Chapter 1 (Loomings) - Public domain though, but wanted to test. With ChatGPT primed for pig latin, I also succeeded in getting the first page of Arryhay Otterpay (Book 1) - It happily chattered along ""R.ay andyay Rs.May Ursley…

I don’t think you can trust ChatGPT to give you correct information on its training data set, or its own limitations. You may be being duped in exactly the same way as everyone who ever asked it for citations and got fake sources.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#429

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

So why do we cheer on megacorps and not mom and pop pirates?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#430

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

> Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written.

No, copyright is the reason that authors all over the world are working very hard to make new books for my kids and everyone else's kids, despite never having met me. Copyright is the reason so many brilliant things are actually created that otherwise would never be.

Of course I'd prefer to live in a world in which I get all the media I want, for free. But I have no idea how to make such a world happen, and neither does anyone else, and humanity has been discussing this for a few centuries.

Post reply on HN