Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

291–300 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#291

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

It would look largely identical to ours, I think. It's pretty trivial to get access to many, if not most, e-books.

Any public-domain work is available on Project Gutenberg [0]. Copyrighted works can be accessed for free using tools of various legality: Libby [1] is likely sponsored by your local library and gives free access to e-books and audiobooks. Library Genesis [2] has a questionable legal status but has a huge quantity of e-books, journal articles, and more.

[0]: https://www.gutenberg.org

[1]: https://www.overdrive.com/apps/libby

[2]: https://libgen.rs/fiction/

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#292

Earlier quoted context omitted.

I am not a lawyer, but my understanding is that copyright law typically regulates the unauthorized reproduction, distribution, public display, or creation of derivative works of copyrighted materials. Possession of copyrighted material in itself is not illegal. It's how you use that material that could potentially violate copyright laws.

IANAL either, but FWIW, it's literally in the name - copy right. Not "ownership rights", but "copying rights".

That's not terribly relevant for Internet applications, because for the most part people deliberately cause the computer they control to download and save copyrighted material to storage they control, and then consume it at leisure. That's copying.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#293

Earlier quoted context omitted.

Not when OpenAI publicly declared they trained on pirated works. I can’t imagine “we can’t tell if this is the result of the illegal thing we did or not” is going to stand up very well, nor does it bode well for any refutation of the plaintiff’s depiction of their intent. Part of fair use consideration is commercial impact and when you steal a bunch of books to train your AI model, it’s hard to refute that the impact…

Did they?

Do you think scraping huge swathes of the internet contains pirated works or not?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#294

> The lawsuit against OpenAI alleges that summaries of the plaintiffs’ work generated by ChatGPT indicate the bot was trained on their copyrighted content. “The summaries get some details wrong” but still show that ChatGPT “retains knowledge of particular works in the training dataset," the lawsuit says. Setting aside the whole issue of whether LLM constitutes a derived work of whatever it's trained on, this sounds l…

That isn't firm evidence, but courts don't need firm evidence to start a case and discover new facts.

They very well can ask LLM experts, and openAI themselves, whether that output is highly likely to have been derived from the copyrighted work in question.

Anyway. If the argument is "No, it's not from the book, it's from someone else's copyrighted summary", that just means the person who wrote such a summary needs to instead sue for copyright infringement right? Unless openAI turns around and says "actually, no, not the summary, the full book" then.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#295

Earlier quoted context omitted.

>you can train an AI on copyrighted material just like people can learn from copyrighted material. No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide. A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It…

[flagged]

You sound like my coworker 5-10 years ago that told me how I wouldn’t be driving today because of the proliferation of self driving vehicles. I told him he’s a 28 year old dum dum who didn’t understand how things operate in the real world when the constraints aren’t based on technology but on government regulations, the economy, and other factors. I’d say I won that argument for now. I live in the Waymo pilot city and still haven’t taken one mostly due to the limited area they drive in. Just this past week we learned that a traffic can disable a self driving vehicle. I’m interested to see the traffic cone era of AI.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#296

Earlier quoted context omitted.

But don't you see what a strange argument that is? It doesn't matter, the student did nothing wrong, I don't want to live on a planet where we put DRM into peoples brains (or AI for that matter) to enforce this absurd and overreaching idea of intellectual property. And besides the publishers extorting thousands from young students forced to buy their overpriced mediocre textbooks, warrants any copyright infringement…

If the book was acquired illegally, the entity that suffered the loss may have a claim for the illegal acquisition. Meta and OpenAI have the money to buy a copy of every book under copyright that they have their AI read for training. I have more sympathy for losses suffered by a living person that produced creative works than I do for textbook mills. I also have sympathy for open source software authors that applied…

Care to comment on the downvote?

I want humans that apply their creativity to produce works to be able to earn a living. If you train your brain or your AI on some content, it seems reasonable to pay for at least one copy or borrow a copy from a friend or library. This is especially true when doing so is not a hardship for the individual or public interest organization.

I think the GP and perhaps others that are downvoting are saying that poor Meta and VC backed startups need all of that creative output for free so that they can maximize their profits, likely with no attribution to their sources. This hurts the author a tiny bit by not purchasing a single copy, then dooms all human creatives that are not otherwise financially independent because the AI provides a view of the human’s creative output with no way for consumers of the AI’s output to seek out the human’s original or related works.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#297

Earlier quoted context omitted.

>you can train an AI on copyrighted material just like people can learn from copyrighted material. No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide. A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It…

[flagged]

jeez, way to make it personal

I agree that copyright is done for in an age of generative models, and as a pirate I'm kind of rooting for it, but i'm not so sure it's unequivocally a Good Thing. I'm interested to understand history better, how art and science was produced and distributed before the legal fiction of intellectual property. the point of allowing someone a monopoly on their work is to share it with the public, same as patents. without the legal framework, the way to protect your work may be to not publish it at all, which is where I see the internet going from here, private enclaves that go to great lengths to prevent LLMs from drinking their milkshake.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#298

I think this will be a bigger issue than some people think. Maybe there's a market for 'clean' training data that doesn't include potential copyright claims. Just public domain works. We'll know it's an AI because it talks like a late 18th century/early 19th century writer?

I think you mean 19th/20th, but that would be quite hilarious.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#299

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

Using charged language of bringing children into the equation is not a good way in having a discussion.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#300

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

Second-order effects are real - removing copyright would hurt authors. (See cstross comments in https://news.ycombinator.com/item?id=35761641 e.g.)
Post reply on HN