Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

131–140 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#131
post #114

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

This is the "guns don't kill people, people do" argument. Not saying that that proves things one way or another, just that assigning responsibility is not really a cut and dry question and many people think that it's important to look prior to the final interaction IANAL but I believe in the US tools that are designed to circumvent copyright are illegal, which makes sense to me inasmuch as one believes that copyright…

Except these models are not designed to circumvent copyright, in fact their primary purpose is to generate non-copyright and non-copyrightable output. It can be induced to product facsimiles of copyright material, but that’s explicitly not the purpose or intent and requires positive action on the users behalf to occur.

Re: Japan’s government will not enforce copyrights on data used in AI training

#132
post #89

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

I strongly agree with this. There's a distinction between "learning from" and "copying". "Learning from" is a transformative process that distills from the observation. This distillation can be as simple as indexing for a search engine, or as complex as a deep neural network. Simply because a neural network can create something that is a copyright violation doesn't mean the training process itself it. A human can see…

> There's a distinction between "learning from" and "copying".

Neural nets can memorize their training data. Generally that isn't what you want, and you strive to eliminate it. However, it could instead be encouraged to happen if someone wanted to exploit this law in order to abuse copyrights.

Re: Japan’s government will not enforce copyrights on data used in AI training

#133
post #22

Earlier quoted context omitted.

There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.

> There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. A few notable differences: 1. Scale: a single art student can't view millions of works in a week. 2. Duplication: a single art student's brain can't be cloned or downloaded into another art student's brain. 3. Speed: a single art student cannot draw or paint thousands of images in a…

This just means students better grow bigger brains!

Re: Japan’s government will not enforce copyrights on data used in AI training

#134

Earlier quoted context omitted.

Training a model isn’t making a copy for your own use, it’s not making a copy at all. It’s converting the original media into a statistical aggregate combined with a lot of other stuff. There’s no copy of the original, even if it’s able to produce a similar product to the original. That’s the specific thing - the aggregation and the lack of direct reproduction in any form is fundamentally not reproducing or copying t…

Copying into RAM during training is making a copy, and can be a copyright violation. https://en.wikipedia.org/wiki/MAI_Systems_Corp._v._Peak_Comp... . However, it seems that there is a later case in the 2nd circuit: https://en.wikipedia.org/wiki/Cartoon_Network,_LP_v._CSC_Hol... .

MAI v. Peak was obviously wrong. It would mean whenever you use someone else's computer, and run licensed software, you're committing copyright infringement. The decision split hairs distinguishing between the current user and the licensee for purposes of legality of making transient copies in memory as part of running the program.

Peak was a repair business. MAI built computers (as in assembled/integrated; I think they were PCs) and had packaged an OS and some software presumably written or modified in-house along with the computer. MAI serviced the whole thing as a unit. So did Peak. MAI sued Peak for copyright infringement because Peak was taking computer repair/maintenance business away from MAI, under the theory that Peak employees operating their clients' MAI computers and software was copyright infringement. (There were other allegations of Peak having unlicensed copies of MAI's software internally, but that's not central to the lawsuit.)

If you have a piece of IP to use to train an IP model with, and you have legal right of access to use that piece of IP (for private purposes), MAI v. Peak doesn't cleanly apply.

MAI v. Peak is also 9th circuit only, and even without the poor reasoning, it should automatically be in doubt because the 9th circuit is notoriously friendly to IP interests, given that it covers Los Angeles.

Re: Japan’s government will not enforce copyrights on data used in AI training

#135

Japan also ranks 3rd (behind the USA & India, with larger populations) in ChatGPT usage: https://www.demandsage.com/chatgpt-statistics/ There's also been discussion of their government using ChatGPT to reduce red tape: https://www.bloomberg.com/news/articles/2023-04-18/japan-gov... It's cool to see Japan and Japanese culture taking techno-optimist stances on AI.

I strongly disagree. They need to actually address problems. Not throw tools at it. The problem Japan seems to have is they don’t understand AI and than they don’t understand software, which is they don’t understand a lot of modern tech. They’ll pay for this mistake just as they paid for being bad at software.

[flagged]

Re: Japan’s government will not enforce copyrights on data used in AI training

#136

What's surprising here is that Japan is usually crazy gung ho on copyright enforcement ... at least against individuals. So it's kind of disgusting to see this relaxation, when it suits some corporate or national interests. https://en.wikipedia.org/wiki/File_sharing_in_Japan "Unlike most other countries, filesharing copyrighted content is not just a civil offense, but a criminal one, with penalties of up to ten years…

I see the other way around, their usual crazy gung ho on copyright enforcement against individuals is disgusting.

This new stance is marvellous and examplifies how free people should interact.

Re: Japan’s government will not enforce copyrights on data used in AI training

#137
post #61

Earlier quoted context omitted.

Superhero costume using logo. Logo inspired by the strongest gem's typical cut with first a monogram the first letter of the name for maximum size and visual clarity. Literally if you asked someone for a recognizable outline of the strongest gemstone's iconic cut you'd get the outline and the rest is an obvious path. Humans might unconsciously, or even by choice, avoid something too similar to something they already…

Unfortunately, I don't believe any of that matters with trademarks. If someone came up with the Superman logo on their own, and released a product that used it, they could not say "but it's a really simple logo" and get a free pass. I'm not sure what that means for ChatGPT, but it would certainly factor into your use of images produced by ChatGPT.

I feel like this is a very strong point that just gets hand-waved away. There are numerous cases where AI-generated content is an exact copy of a derived work. This happens with text, music, and art.

If a we have ai-powered content generation in a video game, and you put into a prompt, "generate 300 mickey mouses, then play some music that sounds like Taylor Swift's new album", and the results look exactly like mickey mouse and the music is Taylor Swift's, it's really difficult to argue that's not copyright infringement.

Yet, people get away with thinking that's not copyright infringement because "the algorithm learned it, like a real human". If the prompt just created a human-designed model, then that is copyright infringement.

The solution might be big corporations create an adversarial network that you can train against to purge copyrighted works from your network.

Re: Japan’s government will not enforce copyrights on data used in AI training

#138

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

> The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use.

Lossy or not, the training data provides value. If all the various someones had not spent time making all the stuff that ends up as training data, then the model it trains would not exist.

If you are going to use someone else's work in order to make something that you are going to profit off of, I believe that original author should be compensated. And should also be able to decide they don't want their work used in that way.

Note that I'm not talking about what existing copyright law says; I'm talking about how I believe we should be regulating this new facet of the industry.

> Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or making it reproduce books etc etc, then distributing the output, would be where the copyright violation occurs.

How is the end-user supposed to know this? Do we seriously believe that everyone who uses generative AI is going to run the output through some sort of process (assuming one even exists) to ensure it's not a substantial enough copy of something some copyrighted work? I certainly don't think this is going to happen.

Regardless, copyright is about distribution. If the a model trained on copyrighted material is considered a copy or derived work of the original work, then distributing that model is, in fact, copyright infringement (absent a successful fair use defense). I'm not saying that's the case, or how a court would look at it, but that's something to consider.

Re: Japan’s government will not enforce copyrights on data used in AI training

#139

Earlier quoted context omitted.

I think liability lies with the person who uses the product to violate copyright. The hosting / producing company didn’t violate copyright if I use their model to make Mickey Mouse pictures. I did.

How can you be certain that the content being generated is non-infringing?

You can’t, I think Fair Use is a fundamentally subjective judgement of a combination of how transformative the work is and the intent and impact of it being distributed.

Re: Japan’s government will not enforce copyrights on data used in AI training

#140
post #89

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

I strongly agree with this. There's a distinction between "learning from" and "copying". "Learning from" is a transformative process that distills from the observation. This distillation can be as simple as indexing for a search engine, or as complex as a deep neural network. Simply because a neural network can create something that is a copyright violation doesn't mean the training process itself it. A human can see…

> A human can see a advertisement for a Marvel movie and then reproduce the Marvel logo. Redistributing (and possibly actually doing that reproduction) that logo is a copyright violation, but the learning process isn't.

I don't think that's correct. That might be trademark infringement, if the logo is a registered trademark, but "seeing something and then drawing it" is in general not copyright infringement.

Post reply on HN