Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

121–130 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#121

Earlier quoted context omitted.

This sounds like the exact opposite of what the title is claiming.

Yeah. Copyright laws in Japan are extremely strict. I would not be surprised at a complete ban on this. This is just a fluff piece like a lot of ads coming in recently.

Copyright law was explicitly changed in 2019 to support AI development. https://storialaw.jp/en/service/bigdata/bigdata-12 https://japannews.yomiuri.co.jp/society/general-news/2023042...

Japan was behind about developing search engine service. Some people argued that it's due to copyright law (there's no fair use) and it shouldn't be repeated again. So the gov want to encourage AI developing by explicit law.

Re: Japan’s government will not enforce copyrights on data used in AI training

#122

Earlier quoted context omitted.

What's the difference between that and someone else who just happens to sound like Eminem in terms out output? As long as you don't market yourself as Eminem that should be completely legal.

I suppose so. I'm just imagining these "ghost AI artists" who publish catalogues of music using the audible likeness of more prolific artists. I know that you could have always just hired an Eminem impersonator and have them lay down tracks...but this technology lets you achieve speed and scale. At least the Eminem impersonator was a real person. This is just a model learned off an artists voice.

These "ghost artists" sound verbatim is the issue I have. A Drake song recently went viral and had tons of hits only to turn out it was made with AI. Fans could not even tell the difference and I most certainly could not either.

Re: Japan’s government will not enforce copyrights on data used in AI training

#123
post #80
post #22

Earlier quoted context omitted.

There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.

Conflating training a model with human learning is wrong. When training a model you are deriving a function that takes some input and produces an output. The issue with copyright and licensing here is that a copy is made and reproduced numerous times when training. The model is not walking around a museum where it is an authorized viewing. It is not a being learning a skill. It is a function. The further issue is tha…

I don't fundamentally disagree with you, but what you are saying doesn't hold water.

> a copy is made and reproduced numerous times when training.

Casually browsing the web creates millions of copies of what are likely the same images and text that models are trained on. Computers cannot move information, they can only copy it and delete the original. Splitting hairs over the semantics of what it means to "copy" isn't a strong argument.

> where it is an authorized viewing

What exactly is an unauthorized viewing of a publicly accessible piece of content online that has been hyperlinked to? If we assume things like robots.txt are respected, what makes the access of that data improper?

> it may output material that competes with the original

An art student could create a forgery. I could craft for myself a replica of a luxury bag. But that's not a crime unless it's done with the intention of deceiving someone or profiting from the work. Intent, after all, is nine tenths of the law.

It's an important right that you should be able to do and create things, even if the sale or distribution of the outputs of those things are prohibited. The ability for a model to produce content which couldn't be distributed shouldn't preempt its existence.

> So you may have copyright violation in distribution of the dataset or a model's output

And neither of those things are the act of training or distributing the model itself!

Re: Japan’s government will not enforce copyrights on data used in AI training

#124
post #89

Earlier quoted context omitted.

I strongly agree with this. There's a distinction between "learning from" and "copying". "Learning from" is a transformative process that distills from the observation. This distillation can be as simple as indexing for a search engine, or as complex as a deep neural network. Simply because a neural network can create something that is a copyright violation doesn't mean the training process itself it. A human can see…

> A human can see a advertisement for a Marvel movie and then reproduce the Marvel logo. Redistributing (and possibly actually doing that reproduction) that logo is a copyright violation, but the learning process isn't. This then becomes about where the liability of that violation lies, and how attractive that is to companies. A human "learning" the marvel logo and reproducing it is violation. How does OpenAPI fit in…

The liability would lay with the company using the LLM product. This could mean that many companies won’t want to take on the risk unless there is decent tooling around warnings of infringement and listing sources.

Re: Japan’s government will not enforce copyrights on data used in AI training

#125

What happens if you ask chatgpt to write you a 7 book harry potter story?

It refuses.

In general, ChatGPT doesn't have copyright books in its training data. But for something as popular as Harry Potter, it's likely plastered all over the internet enough that at least sections of it are in its training data, if not the whole series.

Re: Japan’s government will not enforce copyrights on data used in AI training

#126
post #88
post #77

Earlier quoted context omitted.

> AI models will make 1:1 copies of training data where artists [...] In general I don't think this is the case, assuming you mean generations output from popular text-to-image models. (edit: replied before their comment was edited to include the part on text generation models) For DALL-E 2: I've never seen anyone able to provide a link of supposed copying. Even if you specifically ask it for some prominent work, you…

The more degrees of freedom the less likely independent creation rather than copping occurred. LLM’s recreating training material causes real issues such as Google’s dealing with PII leaks: https://ai.googleblog.com/2020/12/privacy-considerations-in-... If one prompts the GPT-2 language model with the prefix “East Stroudsburg Stroudsburg...”, it will autocomplete a long block of text that contains the full name, phon…

Privacy, where there's a problem if some original data can be inferred/made out (even using a white box attack against the model), is a higher bar than whether an image generator avoids copyright-infringing output under non-adversarial usage. Additionally, compared to image data, text is more prone to exact matches due to lower dimensionality and usually training with less data per parameter.

While it's still a topic deserving of research and mitigation, by the time your information has been scooped up by Common Crawl and trained on by some LLM it's probably in many other places that attackers are more realistically likely to look (search engine caches, Common Crawl downloads, sites specifically for scooping credentials, ...) before trying to extract it from the LLM.

Re: Japan’s government will not enforce copyrights on data used in AI training

#128

Earlier quoted context omitted.

> A human can see a advertisement for a Marvel movie and then reproduce the Marvel logo. Redistributing (and possibly actually doing that reproduction) that logo is a copyright violation, but the learning process isn't. This then becomes about where the liability of that violation lies, and how attractive that is to companies. A human "learning" the marvel logo and reproducing it is violation. How does OpenAPI fit in…

The liability would lay with the company using the LLM product. This could mean that many companies won’t want to take on the risk unless there is decent tooling around warnings of infringement and listing sources.

I think liability lies with the person who uses the product to violate copyright. The hosting / producing company didn’t violate copyright if I use their model to make Mickey Mouse pictures. I did.

Re: Japan’s government will not enforce copyrights on data used in AI training

#129

Earlier quoted context omitted.

Training a model isn’t making a copy for your own use, it’s not making a copy at all. It’s converting the original media into a statistical aggregate combined with a lot of other stuff. There’s no copy of the original, even if it’s able to produce a similar product to the original. That’s the specific thing - the aggregation and the lack of direct reproduction in any form is fundamentally not reproducing or copying t…

Copying into RAM during training is making a copy, and can be a copyright violation. https://en.wikipedia.org/wiki/MAI_Systems_Corp._v._Peak_Comp... . However, it seems that there is a later case in the 2nd circuit: https://en.wikipedia.org/wiki/Cartoon_Network,_LP_v._CSC_Hol... .

How do search engines exist? The internet archive? Caching of image results? Web browser caches? CDNs?

Re: Japan’s government will not enforce copyrights on data used in AI training

#130

Earlier quoted context omitted.

The liability would lay with the company using the LLM product. This could mean that many companies won’t want to take on the risk unless there is decent tooling around warnings of infringement and listing sources.

I think liability lies with the person who uses the product to violate copyright. The hosting / producing company didn’t violate copyright if I use their model to make Mickey Mouse pictures. I did.

How can you be certain that the content being generated is non-infringing?
Post reply on HN