How would that effect, for example, maintaining a large corpus of pirated work for training AI on? Presume that you didn't break any other laws in the process of pirating the works? Could I license a large corpus of pirated work to people explicitly using it for AI training? Like if I downloaded all the books, and then sold access to that collection to someone for AI training, with them agreeing to a EULA preventing…
My understanding in general with piracy is that it goes after those that do the distribution. So you downloading all possible material from where ever as long as it is not illegal is likely rather low risk/penalty. However system will absolutely destroy you if you try to sell or distribute copies of material you collected...
Japan Goes All In: Copyright Doesn't Apply to AI Training
171–180 of 183 posts
Re: Japan Goes All In: Copyright Doesn't Apply to AI Training
#172There is a distinction to be made between inputs and outputs when it comes to AI and copyright. Most people focus on the former, and discuss whether you can train an LLM on copyrighted works or not, but ultimately the issues really only manifest in the latter. An AI training itself on a million newspaper articles can be declared legal, sure, but what happens when it also starts spitting out the same articles with nea…
> This is what the crux of NYT's lawsuit is about, and making laws about AI training isn't going to make a difference to that. The claims of NYT are more than just about training. They're claiming that as part of the ChatGPT software; it _looks up_ stuff in a database of articles. That is beyond _training_. If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think…
I don't see how it is obvious, nor how it is infringement.
How come this doesn't apply to the person who memorized the NYT article, and recited it verbatim?
It's obviously infringement for the person who pressed the "generate" button to produce the article. That's no different from someone who copied a picture using photoshop. However, photoshop itself (and the making of it) does not constitute any infringement whatsoever, as long as at the time of making the application, the sources used are not infringing (which, i presume openAI had the right to view the articles at the time of training).
The crux, to me, is that the information extracted and produced (aka, the neural weights) do not itself constitute any infringement. Using those weights to generate copyrighted stuff is an infringement, but only for the person _doing_ the generation, not on the authors of the weights.
Re: Japan Goes All In: Copyright Doesn't Apply to AI Training
#173Earlier quoted context omitted.
if you create derivitave works of my projects to use in your commercial AI, you're still abusing my work for your own gain.
Copying a style is not a derivative work. You can copy someone’s style just fine. You cannot own a style. I’m perfectly within my rights to make art with your artistic style. And I’m perfectly within my rights to train a bot on my own works. So what’s your argument? What am I not allowed to do with this poorly conceived law?
Re: Japan Goes All In: Copyright Doesn't Apply to AI Training
#174Earlier quoted context omitted.
Copying a style is not a derivative work. You can copy someone’s style just fine. You cannot own a style. I’m perfectly within my rights to make art with your artistic style. And I’m perfectly within my rights to train a bot on my own works. So what’s your argument? What am I not allowed to do with this poorly conceived law?
Ya if you make art to train sure because you own the rights to that art. If you take other peoples art without license as has largely happened, no.
Re: Japan Goes All In: Copyright Doesn't Apply to AI Training
#175So I can see the logic in treating the inputs to the AI training data sets the same way we treat humans learning something. A potential downside is that AI systems can 'mechanise' the creation of material that potentially infringes copyright (in the same way that human generated content can infringe) But a potential upside is that we can 'mechanise' the process by which we judge whether new content infringes the copy…
The issue is the output, not the input. If LLMs were just learning from their training set and generating novel output, the same way a human might, then there'd be no problem. The issue is that generative AI - both for images as well as text - in effect memorize sources as well as learn from from, and can end up regenerating training sources verbatim (or with minimal changes in case of images). I don't think any US c…
I believe human artists consciously adjust their output to avoid copying previous artists too closely. And sometimes they choose to copy very closely or exactly.
Obviously the same feature can be implemented as an option on generative AI systems.
Re: Japan Goes All In: Copyright Doesn't Apply to AI Training
#176Earlier quoted context omitted.
Ya if you make art to train sure because you own the rights to that art. If you take other peoples art without license as has largely happened, no.
Both of these stories involve using the art to the same degree. It’s just a matter of whether you needed a human in the middle
"the art" -> the art you made that you have rights to use, or "the art" that someone else made that you don't have rights to use?
>It’s just a matter of whether you needed a human in the middle
No, your story was you made totally new art that you had rights to use to train your AI. It's not a human in the middle, its a human author who allows you to use the art at all. If the "human in the middle" in round two didn't give you the rights you couldn't use those either. The human is doing the authorship of the work and also allowing or not allowing you to use it. They aren't in the middle, they're 100% of the issue and the difference between allowed and not allowed, human authored or not human authored.
Re: Japan Goes All In: Copyright Doesn't Apply to AI Training
#177So wait. If I encode something near-losslessly into a neural net then it's ok? Books and music are now free ?
It's easier just to pay the $10/month spotify subscription.
Re: Japan Goes All In: Copyright Doesn't Apply to AI Training
#178So wait. If I encode something near-losslessly into a neural net then it's ok? Books and music are now free ?
If you can type a whole book into a word processor, can you now distribute it without breaking copyright?
... or have a computer "type in" a book i.e. file copy i.e. OCR scan (or microphone for audio) (or video-record for video) ...
Re: Japan Goes All In: Copyright Doesn't Apply to AI Training
#179Earlier quoted context omitted.
If you can type a whole book into a word processor, can you now distribute it without breaking copyright?
apparently! ... or have a computer "type in" a book i.e. file copy i.e. OCR scan (or microphone for audio) (or video-record for video) ...
It's no different than a photocopier or a VCR when you're using it for that purpose, it's the end result that matters.
Re: Japan Goes All In: Copyright Doesn't Apply to AI Training
#180Earlier quoted context omitted.
Both of these stories involve using the art to the same degree. It’s just a matter of whether you needed a human in the middle
>Both of these stories involve using the art to the same degree "the art" -> the art you made that you have rights to use, or "the art" that someone else made that you don't have rights to use? >It’s just a matter of whether you needed a human in the middle No, your story was you made totally new art that you had rights to use to train your AI. It's not a human in the middle, its a human author who allows you to use…
It’s ok to make a LegitShady bot as long as you can pay someone $30 to make a few works stylistically similar to yours?
Because if that’s true, you’ve just agreed that the value of your creativity is $0. You don’t even get the $30 here. Nobody really wants your specific works, and your creativity will be freely available. In fact it can probably be synthesized with just a human and some simple tooling in the near future / now.
Is that what you want?