Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

341–350 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#341
post #282

Earlier quoted context omitted.

This isn't completely new, similar issues came up with search engines and this may be seen as 'transformative'. But there may be issues with models that happily reproduce copyrighted texts in their entirety along with other novel issues like models that hallucinate defamatory things or other such problems. Still, I doubt this particular genie can be stuffed back into the bottle, so we'll probably see a lot of litigat…

I agree it's not an entirely new issue. But it's a little different from search results. Say I use the generative paint brush in photoshop. It reproduces a portion of the copyrighted work. I then use the image on an advertising campaign, other merchandise, or post the final product as my own work. Would I be responsible? Would Adobe? Given that retraining these models is not simple, or cheap, would this be just 'cost…

[deleted]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#342

> The lawsuit against OpenAI alleges that summaries of the plaintiffs’ work generated by ChatGPT indicate the bot was trained on their copyrighted content. “The summaries get some details wrong” but still show that ChatGPT “retains knowledge of particular works in the training dataset," the lawsuit says. Setting aside the whole issue of whether LLM constitutes a derived work of whatever it's trained on, this sounds l…

There's an interesting nuance here if you were to put a human in the place of the LLM. We have read thousands of works; does that mean anything we write is derivative?

Humans are special and can create new copyrights. The process of a human brain synthesizing stuff does act as a barrier to copyright infringement.

Machines and algorithms are not legally recognized as being able to author original non-derivative works.

> put a human in the place of the LLM

But also, no, if you have a team of humans doing rote matrix multiplication instead of an LLM, that does not make it so the matrix multiplication removes copyright. Also, at this point LLMs require so much math that you can't replace them with humans, even if the humans have quite fast fingers and calculators.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#343

Earlier quoted context omitted.

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

It would look largely identical to ours, I think. It's pretty trivial to get access to many, if not most, e-books. Any public-domain work is available on Project Gutenberg [0]. Copyrighted works can be accessed for free using tools of various legality: Libby [1] is likely sponsored by your local library and gives free access to e-books and audiobooks. Library Genesis [2] has a questionable legal status but has a huge…

It’s important to note that only some copyrighted works can be accessed for free using the legal options, I’m a member of probably a dozen overdrive supporting libraries and still frequently find titles unavailable for loan of any kind.

I’d love to see an analysis of what % of books are available via libraries around the globe.

Also, the whole DRM thing is a massive pain, audiobooks especially are terrible at allowing side-loading onto a consumer friendly device (such as an MP3 player).

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#344

AI is transformative. Period.

That's a very strong claim that for which you should probably provide some evidence. I just asked ChatGPT to produce a script of a scene from a movie by asking it first to change a single line (which wouldn't be transformative) and then asking it to restore the line to the original. It obliged. Sure, it's probably not the same exact script, but it's not transformative at all. In any case, the issue here isn't whether…

Okay, so you used a tool to duplicate a copyrighted work. You could do the same thing with a word processor. YOU are obviously the one who violated copyright by using the tool that way. I don't understand how a reasonable person could have a different interpretation.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#345

Earlier quoted context omitted.

>you can train an AI on copyrighted material just like people can learn from copyrighted material. No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide. A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It…

Derivative works are typically things like language translations or film adaptions of a novel. A large language model is something like a probability breakdown of the order of word fragments in a body of text. It's a collection of statistics and math. It's different. Now, can you get it to output a derivative work? Maybe. Is every output a derivative work? Maybe not.

Things like sports statistics and directories of phone numbers are not copyrightable. Maybe models fall into a similar category. Could be. It’ll be interesting to see how this shakes out over the next few years.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#346

Earlier quoted context omitted.

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

There are already many more public domain books than people are inclined to read: https://www.gutenberg.org/ https://librivox.org/ many of which form the basis for an education: https://news.ycombinator.com/item?id=34630153 And most of which, when in copyright, paid their authors quite handsomely in terms of royalties. If you believe that books should exist without copyright, then one has to ask --- how many books ha…

I’d more ask what the cost vs benefits are of keeping the existing scheme, it’s not free to run all these DRM services, prosecute offenders etc…

Not to say I support no copyright…

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#347
post #213

Earlier quoted context omitted.

> What if it turns out that OpenAI bought a copy of every book ingested be ChatGPT? That still doesn't necessarily confer to them the right to use it to train a model and generate derivative works based on purchased content.

What if that turns out to be completely irrelevant? Let's say, for the sake of argument, that I knew absolutely nothing about contract law and was then filmed stealing a book you wrote on the subject from a book store. I then started a business where I would answer questions about contract law, based solely on what I learned from the book. Of course, my memory isn't perfect, but I don't like to admit when I'm wrong,…

Replace "answer questions about contract law, based solely on what I learned from the book" with "generate cartoon images, based solely on what ML learned from Disney IP" and see how badly that will go.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#349

Earlier quoted context omitted.

There are already many more public domain books than people are inclined to read: https://www.gutenberg.org/ https://librivox.org/ many of which form the basis for an education: https://news.ycombinator.com/item?id=34630153 And most of which, when in copyright, paid their authors quite handsomely in terms of royalties. If you believe that books should exist without copyright, then one has to ask --- how many books ha…

>If you believe that books should exist without copyright, then one has to ask --- how many books have you written which you have explicitly placed in the public domain? Or, how many authors have you patronized so as to fund their writing so that they can publish their works freely? Lol Edit: I should probably clarify here. While I can’t speak for OP, I can say that, for some reason, I am sure there are people who ha…

I've authored quite a bit of book-like content which has been made freely available:

- the Shapeoko wiki (still available on archive.org)

- a couple of articles for TUGboat

- edited a couple of texts on wikibooks trying to make them better

- currently working on https://willadams.gitbook.io/design-into-3d/

The Venn diagram of folks who don't believe in copyright and those who have actually produced something other folks want to read is quite sparse, excepting the odd manifesto.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#350
post #131

Earlier quoted context omitted.

Not when OpenAI publicly declared they trained on pirated works. I can’t imagine “we can’t tell if this is the result of the illegal thing we did or not” is going to stand up very well, nor does it bode well for any refutation of the plaintiff’s depiction of their intent. Part of fair use consideration is commercial impact and when you steal a bunch of books to train your AI model, it’s hard to refute that the impact…

Please read more carefully. OpenAI never “declared they trained on pirated works.”

However, the root of this thread quotes Meta admitting exactly that.
Post reply on HN