Earlier quoted context omitted.
This isn't completely new, similar issues came up with search engines and this may be seen as 'transformative'. But there may be issues with models that happily reproduce copyrighted texts in their entirety along with other novel issues like models that hallucinate defamatory things or other such problems. Still, I doubt this particular genie can be stuffed back into the bottle, so we'll probably see a lot of litigat…
I agree it's not an entirely new issue. But it's a little different from search results. Say I use the generative paint brush in photoshop. It reproduces a portion of the copyrighted work. I then use the image on an advertising campaign, other merchandise, or post the final product as my own work. Would I be responsible? Would Adobe? Given that retraining these models is not simple, or cheap, would this be just 'cost…
Sarah Silverman is suing OpenAI and Meta for copyright infringement
341–350 of 599 posts
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#342> The lawsuit against OpenAI alleges that summaries of the plaintiffs’ work generated by ChatGPT indicate the bot was trained on their copyrighted content. “The summaries get some details wrong” but still show that ChatGPT “retains knowledge of particular works in the training dataset," the lawsuit says. Setting aside the whole issue of whether LLM constitutes a derived work of whatever it's trained on, this sounds l…
There's an interesting nuance here if you were to put a human in the place of the LLM. We have read thousands of works; does that mean anything we write is derivative?
Machines and algorithms are not legally recognized as being able to author original non-derivative works.
> put a human in the place of the LLM
But also, no, if you have a team of humans doing rote matrix multiplication instead of an LLM, that does not make it so the matrix multiplication removes copyright. Also, at this point LLMs require so much math that you can't replace them with humans, even if the humans have quite fast fingers and calculators.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#343Earlier quoted context omitted.
Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…
It would look largely identical to ours, I think. It's pretty trivial to get access to many, if not most, e-books. Any public-domain work is available on Project Gutenberg [0]. Copyrighted works can be accessed for free using tools of various legality: Libby [1] is likely sponsored by your local library and gives free access to e-books and audiobooks. Library Genesis [2] has a questionable legal status but has a huge…
I’d love to see an analysis of what % of books are available via libraries around the globe.
Also, the whole DRM thing is a massive pain, audiobooks especially are terrible at allowing side-loading onto a consumer friendly device (such as an MP3 player).
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#344AI is transformative. Period.
That's a very strong claim that for which you should probably provide some evidence. I just asked ChatGPT to produce a script of a scene from a movie by asking it first to change a single line (which wouldn't be transformative) and then asking it to restore the line to the original. It obliged. Sure, it's probably not the same exact script, but it's not transformative at all. In any case, the issue here isn't whether…
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#345Earlier quoted context omitted.
>you can train an AI on copyrighted material just like people can learn from copyrighted material. No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide. A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It…
Derivative works are typically things like language translations or film adaptions of a novel. A large language model is something like a probability breakdown of the order of word fragments in a body of text. It's a collection of statistics and math. It's different. Now, can you get it to output a derivative work? Maybe. Is every output a derivative work? Maybe not.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#346Earlier quoted context omitted.
Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…
There are already many more public domain books than people are inclined to read: https://www.gutenberg.org/ https://librivox.org/ many of which form the basis for an education: https://news.ycombinator.com/item?id=34630153 And most of which, when in copyright, paid their authors quite handsomely in terms of royalties. If you believe that books should exist without copyright, then one has to ask --- how many books ha…
Not to say I support no copyright…
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#347Earlier quoted context omitted.
> What if it turns out that OpenAI bought a copy of every book ingested be ChatGPT? That still doesn't necessarily confer to them the right to use it to train a model and generate derivative works based on purchased content.
What if that turns out to be completely irrelevant? Let's say, for the sake of argument, that I knew absolutely nothing about contract law and was then filmed stealing a book you wrote on the subject from a book store. I then started a business where I would answer questions about contract law, based solely on what I learned from the book. Of course, my memory isn't perfect, but I don't like to admit when I'm wrong,…
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#348AI is transformative. Period.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#349Earlier quoted context omitted.
There are already many more public domain books than people are inclined to read: https://www.gutenberg.org/ https://librivox.org/ many of which form the basis for an education: https://news.ycombinator.com/item?id=34630153 And most of which, when in copyright, paid their authors quite handsomely in terms of royalties. If you believe that books should exist without copyright, then one has to ask --- how many books ha…
>If you believe that books should exist without copyright, then one has to ask --- how many books have you written which you have explicitly placed in the public domain? Or, how many authors have you patronized so as to fund their writing so that they can publish their works freely? Lol Edit: I should probably clarify here. While I can’t speak for OP, I can say that, for some reason, I am sure there are people who ha…
- the Shapeoko wiki (still available on archive.org)
- a couple of articles for TUGboat
- edited a couple of texts on wikibooks trying to make them better
- currently working on https://willadams.gitbook.io/design-into-3d/
The Venn diagram of folks who don't believe in copyright and those who have actually produced something other folks want to read is quite sparse, excepting the odd manifesto.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#350Earlier quoted context omitted.
Not when OpenAI publicly declared they trained on pirated works. I can’t imagine “we can’t tell if this is the result of the illegal thing we did or not” is going to stand up very well, nor does it bode well for any refutation of the plaintiff’s depiction of their intent. Part of fair use consideration is commercial impact and when you steal a bunch of books to train your AI model, it’s hard to refute that the impact…
Please read more carefully. OpenAI never “declared they trained on pirated works.”