I understand that nobody wants a multi billion dollar company to take their works without giving compensation, but I worry all of this could move dangerously close to allowing concepts and styles to be copyrighted and crack down on transformative works.
Sarah Silverman is suing OpenAI and Meta for copyright infringement
411–420 of 599 posts
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#412Could one argue that training AI systems constitutes an educational purpose, invoking the copyright exemption?
If you are making a commercial product I think not. Or do you mean educational as in educating the AI itself?
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#413Earlier quoted context omitted.
You're analyzing the factors with respect to the output of the model, not the weights: > The purpose and character is absolutely heavily commercial and makes a great deal of money for the companies building the AIs. That's assuming that it's Microsoft/OpenAI. Suppose a non-profit trains a model and releases the weights for free. > There's nothing about the works used for AI training that makes them any less entitled…
> You can't find a copy of any specific work anywhere in the weights. You won't find a copy of any specific work anywhere in the compressed form of a file, either, but when you decompress it you find the complete work. And many large AIs can recite, verbatim or near-verbatim, many complete works. Yes, they might get a word wrong, but that doesn't nullify the point that they're trained on the entire work and to a firs…
Sure you will. It's right there, in PNG encoding or what have you. With nothing more than the compressed file and general purpose tools you can reliably put it on your screen.
> And many large AIs can recite, verbatim or near-verbatim, many complete works.
This is not the common case and it's not even clear that the reason it can sometimes do this is having been trained on the complete work. The more likely cause is having seen a large number of excerpts from the work which patch together into the whole thing, because it's more likely to output that text if it has seen it multiple times.
This is also clearly not the intent, purpose or typical use of the model. It has no ability to consistently do that.
> Many of the consumers of image models are in fact generating art that they previously would have commissioned from an artist. ... Consumers of code models are, in fact, potentially reducing the total demand for novice programmers.
Those are different people than the holder of the copyright on an arbitrary piece of the training data. Is a fair use determination supposed care about the effect on the market for some entirely independent work from an unrelated third party?
> The model derived from a pile of artistic works is, in fact, directly competing with those artistic works.
It's indirectly competing with artistic works in general.
I mean take the computer out of the loop and think about it this way. Art teachers everywhere reproduce a bunch of existing works for classroom use to train new artists who go into competition with the original artists. That is obviously going to indirectly impact the market for artistic works in general, but I don't think that's how that factor works.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#414Earlier quoted context omitted.
In that particular case if you have an enterprise licence, Adobe have accepted responsibility: >Adobe is so confident its Firefly generative AI won’t breach copyright that it’ll cover your legal bills The offer is available only to users of its enterprise Firefly product, which launches today https://www.fastcompany.com/90906560/adobe-feels-so-confiden...
If they are that confident in their product why do they only indemnify users of their enterprise Firefly product? Smells to me like a bad case of corporate marketing bullshit.
It is likely to be the large companies that worry about getting sued and tell their employees not to use the functionality. By offering protection Adobe get to sell more to the people paying millions a year.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#415Earlier quoted context omitted.
How about you go figure out this new economic model, and come back when it's ready. Until then, the existing model will persist, thank you
Mandatory licensing seems fine. The government charges a tax on everyone. Revenue is handed to authors based on how many times the work is used. There are some questions around how to weigh things like societal importance of the work. (“Not melting down nuclear reactors for dummies” seems like it should get more money per view than “poodles in outer space vol XXXII”, despite likely having lower readership)
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#416> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…
Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…
A company that believes in strong intellectual property rights protection is using resources that blatantly ignore intellectual property rights to get access to the content for free.
I agree with you, however, that it's an argument in favor of abolishing strong intellectual property rights. At least for OpenAI's products.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#417Earlier quoted context omitted.
Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…
There are already many more public domain books than people are inclined to read: https://www.gutenberg.org/ https://librivox.org/ many of which form the basis for an education: https://news.ycombinator.com/item?id=34630153 And most of which, when in copyright, paid their authors quite handsomely in terms of royalties. If you believe that books should exist without copyright, then one has to ask --- how many books ha…
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#418Earlier quoted context omitted.
I’d more ask what the cost vs benefits are of keeping the existing scheme, it’s not free to run all these DRM services, prosecute offenders etc… Not to say I support no copyright…
The benefit of the existing scheme is that new works get created, and some of them are even copy-edited and published. How many texts are created which are explicitly placed in the public domain and from which the authors have made a conscious decision not to profit thereby?
Most people who write non-fiction books do it because they want to contribute to human knowledge and be recognized as an expert in a particular field, not because they think that writing a differential geometry textbook is their path to riches. With the internet, more and more text books are made freely available by their authors - the reason this didn’t happen in the past is because there was no other way to pass knowledge around than teaming up with a publisher who is able to put your knowledge on to dead trees. It’s fair game to put older books on the internet, so that the whole world can benefit, not just rich people in rich countries.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#419Earlier quoted context omitted.
IANAL either, but FWIW, it's literally in the name - copy right. Not "ownership rights", but "copying rights".
That's not terribly relevant for Internet applications, because for the most part people deliberately cause the computer they control to download and save copyrighted material to storage they control, and then consume it at leisure. That's copying.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#420If they ripped all of Bibliotik, the more interesting story to me is how they were able to get it all without hitting ratio requirements? Super fast internet that downloaded all they could before being ratio banned, overwhelmingly fast internet that was hopping on all the popular torrents to slowly build up ratio?