I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…
This is the "guns don't kill people, people do" argument. Not saying that that proves things one way or another, just that assigning responsibility is not really a cut and dry question and many people think that it's important to look prior to the final interaction IANAL but I believe in the US tools that are designed to circumvent copyright are illegal, which makes sense to me inasmuch as one believes that copyright…
Japan’s government will not enforce copyrights on data used in AI training
131–140 of 426 posts
Re: Japan’s government will not enforce copyrights on data used in AI training
#132I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…
I strongly agree with this. There's a distinction between "learning from" and "copying". "Learning from" is a transformative process that distills from the observation. This distillation can be as simple as indexing for a search engine, or as complex as a deep neural network. Simply because a neural network can create something that is a copyright violation doesn't mean the training process itself it. A human can see…
Neural nets can memorize their training data. Generally that isn't what you want, and you strive to eliminate it. However, it could instead be encouraged to happen if someone wanted to exploit this law in order to abuse copyrights.
Re: Japan’s government will not enforce copyrights on data used in AI training
#133Earlier quoted context omitted.
There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.
> There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. A few notable differences: 1. Scale: a single art student can't view millions of works in a week. 2. Duplication: a single art student's brain can't be cloned or downloaded into another art student's brain. 3. Speed: a single art student cannot draw or paint thousands of images in a…
Re: Japan’s government will not enforce copyrights on data used in AI training
#134Earlier quoted context omitted.
Training a model isn’t making a copy for your own use, it’s not making a copy at all. It’s converting the original media into a statistical aggregate combined with a lot of other stuff. There’s no copy of the original, even if it’s able to produce a similar product to the original. That’s the specific thing - the aggregation and the lack of direct reproduction in any form is fundamentally not reproducing or copying t…
Copying into RAM during training is making a copy, and can be a copyright violation. https://en.wikipedia.org/wiki/MAI_Systems_Corp._v._Peak_Comp... . However, it seems that there is a later case in the 2nd circuit: https://en.wikipedia.org/wiki/Cartoon_Network,_LP_v._CSC_Hol... .
Peak was a repair business. MAI built computers (as in assembled/integrated; I think they were PCs) and had packaged an OS and some software presumably written or modified in-house along with the computer. MAI serviced the whole thing as a unit. So did Peak. MAI sued Peak for copyright infringement because Peak was taking computer repair/maintenance business away from MAI, under the theory that Peak employees operating their clients' MAI computers and software was copyright infringement. (There were other allegations of Peak having unlicensed copies of MAI's software internally, but that's not central to the lawsuit.)
If you have a piece of IP to use to train an IP model with, and you have legal right of access to use that piece of IP (for private purposes), MAI v. Peak doesn't cleanly apply.
MAI v. Peak is also 9th circuit only, and even without the poor reasoning, it should automatically be in doubt because the 9th circuit is notoriously friendly to IP interests, given that it covers Los Angeles.
Re: Japan’s government will not enforce copyrights on data used in AI training
#135Japan also ranks 3rd (behind the USA & India, with larger populations) in ChatGPT usage: https://www.demandsage.com/chatgpt-statistics/ There's also been discussion of their government using ChatGPT to reduce red tape: https://www.bloomberg.com/news/articles/2023-04-18/japan-gov... It's cool to see Japan and Japanese culture taking techno-optimist stances on AI.
I strongly disagree. They need to actually address problems. Not throw tools at it. The problem Japan seems to have is they don’t understand AI and than they don’t understand software, which is they don’t understand a lot of modern tech. They’ll pay for this mistake just as they paid for being bad at software.
Re: Japan’s government will not enforce copyrights on data used in AI training
#136What's surprising here is that Japan is usually crazy gung ho on copyright enforcement ... at least against individuals. So it's kind of disgusting to see this relaxation, when it suits some corporate or national interests. https://en.wikipedia.org/wiki/File_sharing_in_Japan "Unlike most other countries, filesharing copyrighted content is not just a civil offense, but a criminal one, with penalties of up to ten years…
This new stance is marvellous and examplifies how free people should interact.
Re: Japan’s government will not enforce copyrights on data used in AI training
#137Earlier quoted context omitted.
Superhero costume using logo. Logo inspired by the strongest gem's typical cut with first a monogram the first letter of the name for maximum size and visual clarity. Literally if you asked someone for a recognizable outline of the strongest gemstone's iconic cut you'd get the outline and the rest is an obvious path. Humans might unconsciously, or even by choice, avoid something too similar to something they already…
Unfortunately, I don't believe any of that matters with trademarks. If someone came up with the Superman logo on their own, and released a product that used it, they could not say "but it's a really simple logo" and get a free pass. I'm not sure what that means for ChatGPT, but it would certainly factor into your use of images produced by ChatGPT.
If a we have ai-powered content generation in a video game, and you put into a prompt, "generate 300 mickey mouses, then play some music that sounds like Taylor Swift's new album", and the results look exactly like mickey mouse and the music is Taylor Swift's, it's really difficult to argue that's not copyright infringement.
Yet, people get away with thinking that's not copyright infringement because "the algorithm learned it, like a real human". If the prompt just created a human-designed model, then that is copyright infringement.
The solution might be big corporations create an adversarial network that you can train against to purge copyrighted works from your network.
Re: Japan’s government will not enforce copyrights on data used in AI training
#138I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…
Lossy or not, the training data provides value. If all the various someones had not spent time making all the stuff that ends up as training data, then the model it trains would not exist.
If you are going to use someone else's work in order to make something that you are going to profit off of, I believe that original author should be compensated. And should also be able to decide they don't want their work used in that way.
Note that I'm not talking about what existing copyright law says; I'm talking about how I believe we should be regulating this new facet of the industry.
> Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or making it reproduce books etc etc, then distributing the output, would be where the copyright violation occurs.
How is the end-user supposed to know this? Do we seriously believe that everyone who uses generative AI is going to run the output through some sort of process (assuming one even exists) to ensure it's not a substantial enough copy of something some copyrighted work? I certainly don't think this is going to happen.
Regardless, copyright is about distribution. If the a model trained on copyrighted material is considered a copy or derived work of the original work, then distributing that model is, in fact, copyright infringement (absent a successful fair use defense). I'm not saying that's the case, or how a court would look at it, but that's something to consider.
Re: Japan’s government will not enforce copyrights on data used in AI training
#139Earlier quoted context omitted.
I think liability lies with the person who uses the product to violate copyright. The hosting / producing company didn’t violate copyright if I use their model to make Mickey Mouse pictures. I did.
How can you be certain that the content being generated is non-infringing?
Re: Japan’s government will not enforce copyrights on data used in AI training
#140I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…
I strongly agree with this. There's a distinction between "learning from" and "copying". "Learning from" is a transformative process that distills from the observation. This distillation can be as simple as indexing for a search engine, or as complex as a deep neural network. Simply because a neural network can create something that is a copyright violation doesn't mean the training process itself it. A human can see…
I don't think that's correct. That might be trademark infringement, if the logo is a registered trademark, but "seeing something and then drawing it" is in general not copyright infringement.