Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

301–310 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#301

Earlier quoted context omitted.

I’m imagining a world that looks just about the same as this one does. A larger book library doesn’t automatically make that medium more appealing to kids than what Mr Beast, Unspeakable, and the other crap kids love are doing.

...for the global middle class? Maybe. For the world as a whole? Definite differences. It just seems like a super jaded "kids these days" thing to hate on them for consuming easily accessed, free content- and acting like the global literacy and intellectual capital would remain unaffected.

Also, piracy exists and is perfectly reasonable for an individual, especially when they are not able to afford the book, to use to get access to a book.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#302
post #244

Earlier quoted context omitted.

It is distribution of copyrighted material without permission of the author that is illegal, when you download you're not distributing so it isn't illegal (unless you're using something like BitTorrent that also distributes it while you're downloading it).

In the US permission is required to make copies, prepare derivative works, distribute copies, publicly perform the work, or publicly display the work [1]. [1] https://www.law.cornell.edu/uscode/text/17/106

Right, the question is, when my computer requests a file from your computer, which one of us is "making a copy" ? It becomes less ambiguous to ask who is doing the publishing.

In a physical analogy, if someone is selling bootleg DVDs on the street, I don't think anyone ever got busted for being a customer.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#303

Earlier quoted context omitted.

[flagged]

You sound like my coworker 5-10 years ago that told me how I wouldn’t be driving today because of the proliferation of self driving vehicles. I told him he’s a 28 year old dum dum who didn’t understand how things operate in the real world when the constraints aren’t based on technology but on government regulations, the economy, and other factors. I’d say I won that argument for now. I live in the Waymo pilot city an…

I'm basing this on what US courts have decided. You're free to disagree with them all you'd like, but AI-generated art is not copyrightable. We're seeing an explosion in non-copyrightable art, and when we get down to some small fraction of art being copyrightable, nobody will give a shit about copyright anymore. You also talk about the economy, and guess where the economoc incentives are aligned towards? Hint: it's not towards having expensive humans generate art.

I guess we'll just have to feel that the other is wrong as we wait to see what happens, but I'm betting on existing US regulation and basic human behavior. That's a tall order to bet against.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#304

Earlier quoted context omitted.

I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.

> I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. I would be very happy if either a court or lawmakers decided that copyright itself was unconscionable. That isn't what's going to happen, though. And I think it's incredibly unacceptable if a court…

You're analyzing the factors with respect to the output of the model, not the weights:

> The purpose and character is absolutely heavily commercial and makes a great deal of money for the companies building the AIs.

That's assuming that it's Microsoft/OpenAI. Suppose a non-profit trains a model and releases the weights for free.

> There's nothing about the works used for AI training that makes them any less entitled to copyright protections or more permissive of fair use than anything else. They're not unpublished, they're not merely collections of facts or ideas, etc.

The models aren't trained on works of a particular nature, they're trained on whatever they can find, so this doesn't really mean anything until you're talking about a specific work.

> AI training uses entire works, not excerpts.

The weights don't contain entire works. They contain statistics about entire works, but that's not the same thing. You can't find a copy of any specific work anywhere in the weights. Nobody can give you a piece of code that will decode the weights into all the original works.

> AI models are having a massive effect on the market for and value of the works they train on, as is being widely discussed in multiple industries.

That's not how that factor works (and it's the most important one). No one is buying the model weights so they can read them like a novel. Typically the consumers of the weights are software developers or content creators, whereas the consumers of the original text or image are fans.

To make this a little clearer, suppose the purpose of the model isn't to generate content, it's to generate recommendations. Then the company operating it takes a list of content anyone likes, uses the model to show you a list of all the other content you might like and lets you sort by price. Which makes it easier to find competing content which is available for less money, which increases competition. The incumbents might hate this, and it might even lower their profits, but that's not the kind of effect on the market this factor is supposed to be about.

Moreover, suppose you had a model trained entirely on public domain works. Obviously this can't be copyright infringement even if it's extremely effective at producing new works that compete heavily with works still under copyright. But if you added some specific work still under copyright to the model, it would only be a marginal difference. The effect on the market for that specific work of adding that specific work to the model would be negligible. It's the technology itself that provides the competition, not the accretion of any particular work in the weights.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#305
post #282

I think this will be a bigger issue than some people think. Maybe there's a market for 'clean' training data that doesn't include potential copyright claims. Just public domain works. We'll know it's an AI because it talks like a late 18th century/early 19th century writer?

This isn't completely new, similar issues came up with search engines and this may be seen as 'transformative'. But there may be issues with models that happily reproduce copyrighted texts in their entirety along with other novel issues like models that hallucinate defamatory things or other such problems. Still, I doubt this particular genie can be stuffed back into the bottle, so we'll probably see a lot of litigat…

I agree it's not an entirely new issue. But it's a little different from search results. Say I use the generative paint brush in photoshop. It reproduces a portion of the copyrighted work. I then use the image on an advertising campaign, other merchandise, or post the final product as my own work. Would I be responsible? Would Adobe? Given that retraining these models is not simple, or cheap, would this be just 'cost of doing business?' Would I be able to buy insurance for this?

Enquiring minds want to know.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#306

I think this will be a bigger issue than some people think. Maybe there's a market for 'clean' training data that doesn't include potential copyright claims. Just public domain works. We'll know it's an AI because it talks like a late 18th century/early 19th century writer?

I think you mean 19th/20th, but that would be quite hilarious.

Dear Mr. Smith:

Your employer has generously agreed to offer you a position, quite reasonable, and with many great benefits at the venerable firm of 'Zumba'. Our interest is that you should join our staff forthwith and at the earliest date. A cab and man has been sent to retrieve you and bring you to our offices to sign all the necessary documents. Our offer is for a monthly stipend of five pounds, two shillings, and sixpence to be paid at the end of the month.

Thank you,

Most Humbly,

Hirebot 2347

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#307

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

Then most people stop writing books because they can't get paid for their time/effort and ~every child will be stuck with outdated knowledge within a decade.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#308
post #307

Earlier quoted context omitted.

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

Then most people stop writing books because they can't get paid for their time/effort and ~every child will be stuck with outdated knowledge within a decade.

Or maybe we could figure out a new economic model, instead of blindly sticking with one based on the limitations of the pre-digital age.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#309

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

There are already many more public domain books than people are inclined to read:

https://www.gutenberg.org/

https://librivox.org/

many of which form the basis for an education:

https://news.ycombinator.com/item?id=34630153

And most of which, when in copyright, paid their authors quite handsomely in terms of royalties.

If you believe that books should exist without copyright, then one has to ask --- how many books have you written which you have explicitly placed in the public domain? Or, how many authors have you patronized so as to fund their writing so that they can publish their works freely? Or, if neither of these applies, how do you propose to compensate authors for the efforts and labours of writing?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#310

Earlier quoted context omitted.

I’m imagining a world that looks just about the same as this one does. A larger book library doesn’t automatically make that medium more appealing to kids than what Mr Beast, Unspeakable, and the other crap kids love are doing.

...for the global middle class? Maybe. For the world as a whole? Definite differences. It just seems like a super jaded "kids these days" thing to hate on them for consuming easily accessed, free content- and acting like the global literacy and intellectual capital would remain unaffected.

How many of these kids have read what percentage of the books which are legitimately freely available?

I've never encountered a kid (other than my own) who has read:

https://mathcs.clarku.edu/~djoyce/java/elements/elements.htm...

but have encountered many others who struggle with geometry.

Post reply on HN