Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

541–550 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#541
post #97

Earlier quoted context omitted.

Not necessarily, because the models have an element of randomness. Also, I was under the impression that ChatGPT has more "safeguards" (manifesting as a refusal to answer questions) than the raw API.

I don’t doubt the poster was telling the truth when they said they asked for a summary of the book and didn’t get one. It refutes the idea that chatgpt’s inability to provide a summary means it didn’t scan the original text: since it can provide a summary, the argument is entirely spurious.

The summary it provides is entirely wrong except for the name of the main character.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#542

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

That's not parent's point. Their point seems to be that large companies don't suffer the same consequences from a crime as any layman.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#543

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

What an absolute pile of nonsense. People who author creative works deserve to have control of them and make some money - otherwise you'll soon find we have far fewer great authors, artists, etc.

This is essentially the same as saying builders charging for houses is the problem with the housing market, so we're going to phase out paying builders.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#544

Earlier quoted context omitted.

> If everyone is allowed to steal books Nothing was stolen- just copied.

> Nothing was stolen- just copied. This typical semantic-pedantry line from piracy apologists misses the point - piracy is theft-adjacent even if you get to pick your use of "theft". Incidentally, my definition of "theft", and that of most content creators, includes the act of consuming something without compensating the creator on their terms - which includes piracy.

[flagged]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#545

Earlier quoted context omitted.

It's worth challenging the length of copyright. 20 years seems good enough for high tech R&D, probably good for books as well.

Sure! Yes! I agree! 100 years is way too long. 20 years is much more reasonable. But the comment that I was responding to (and many others in this thread) are advocating for the complete removal of copyright, and that's what I'm responding to.

Most book sales are done within 5 years. It's drips after that.

On top of that, knowledge moves too fast these days for a 20 year right to be useful for society.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#546

Earlier quoted context omitted.

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

What an absolute pile of nonsense. People who author creative works deserve to have control of them and make some money - otherwise you'll soon find we have far fewer great authors, artists, etc. This is essentially the same as saying builders charging for houses is the problem with the housing market, so we're going to phase out paying builders.

> otherwise you'll soon find we have far fewer great authors, artists, etc.

It's been long understood that this idea isn't based on any real evidence. Creators create because they like to create things. Adding money to the mix tends to ruin most creative endeavours. Look up beautification or note how Googles only good search results involve the keyword Reddit.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#547
post #38

>On information and belief, the reason ChatGPT can accurately summarize a certain copyrighted book is because that book was copied by OpenAI and ingested by the underlying OpenAI Language Model (either GPT-3.5 or GPT-4) as part of its training data. While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false. First, even if the mode…

Isn't part of the problem that some of the training data is retained by the model and used during response generation? In that case it's not just that the copyrighted book was used as training data but that some part of the book has been retained by the model. So now my model is using copyrighted material while it runs. Here's an example of a model that retained enough image data to reconstruct a reasonable facsimile of the training image.

https://www.theregister.com/2023/02/06/uh_oh_attackers_can_e...

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#548

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

GP makes no remark on the morality/practicality of copyright. Also, having people sue big companies for copyright might lead to more of what you're arguing for, in a show them the taste of their own medicine way

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#549

Earlier quoted context omitted.

It's worth challenging the length of copyright. 20 years seems good enough for high tech R&D, probably good for books as well.

Sure! Yes! I agree! 100 years is way too long. 20 years is much more reasonable. But the comment that I was responding to (and many others in this thread) are advocating for the complete removal of copyright, and that's what I'm responding to.

100 years is arguably “unconstitutional.”

Constitution days “for a limited time”. Death of author + 99 is effectively unlimited to the perspective of typical human

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#550

Earlier quoted context omitted.

What an absolute pile of nonsense. People who author creative works deserve to have control of them and make some money - otherwise you'll soon find we have far fewer great authors, artists, etc. This is essentially the same as saying builders charging for houses is the problem with the housing market, so we're going to phase out paying builders.

> otherwise you'll soon find we have far fewer great authors, artists, etc. It's been long understood that this idea isn't based on any real evidence. Creators create because they like to create things. Adding money to the mix tends to ruin most creative endeavours. Look up beautification or note how Googles only good search results involve the keyword Reddit.

Theres "creators create because they like to create things" and "creators create things because they want to make a living off of what they create". If artists/authors/musicians/etc aren't going to be paid for what they create, they can't make a living doing it. If they can't make a living doing it, that severely limits their opportunity and time available for creating things since they could only do it as a hobby (unless we bring back royal patronage or something). Many of the best artistic works we have came from people who were able to commit 100% to the creative process. That's gonna be real hard to do if you can't pay for food and housing.
Post reply on HN