Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

161–170 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#161

The original links here, to the actual question asked and the answer by the minister: https://kiitaka.net/21312/

I thought it would be too ironic for people to misunderstand this based on a machine translated version, so here is a genuine, human translation of the transcript, with boring bits redacted.

Kii: Next question, again regarding generative AI. I would like to ask from the two perspectives of copyright protection and educational use. [...] First, can we understand that Japanese law permits the use of works for information analysis, both for non-commercial and commercial purposes, and acts other than copying, and using content that was uploaded illegally?

Nagaoka: Use for non-commercial information analysis is permitted under Article 30-4 of the copyright act, provided that the purpose is not the enjoyment of the ideas and emotions expressed in the copyrighted work.

Kii: Minister, I asked about four aspects of use for information analysis: non-commercial use, commercial use, acts other than copying, and illegally uploaded content. Please address the other three.

Nagaoka: Use for commercial purposes is permitted under Article 30-4 of the copyright act, provided that the purpose is not the enjoyment of the ideas and emotions expressed in the copyrighted work, because that Article does not distinguish between information analysis for commercial or non-commercial purposes.

Regarding copying, Article 30-4 of the Copyright Act does not distinguish based on the method of use, so use by means other than copying is permitted provided that the criteria are met.

[...] Regarding content obtained from piracy sites and the like, [...] illegal uploading itself is infringement of copyright, and is subject to a damage claim, petition for injunction, or criminal punishment. However, it is not practically feasible to identify whether any particular work in a large collection obtained from the internet is copyrighted or not, so making this a criterion for information analysis would make it difficult to use information analysis for Big Data.

In addition, as the use of a work for information analysis is not use for the purpose of enjoyment of the ideas or emotions expressed in the work, and even if [it were used in that manner] it would not overlap with the original market for the use of the work, so it is not considered to harm the interests of the copyright holder that are protected by the Copyright Act.

As such, Article 30-4 of the Copyright Act does not have the legality of the work as a criterion.

Kii: Minister, based on your answer, I think the greatest issue is that there is no protection against use that goes against the intentions of the creator or the copyright holder. I believe that new regulations will be necessary to address this point; will you consider such new regulations?

Nagaoka: Article 30-4 of the Copyright Act provides for use that is not for the purpose of enjoying the ideas or emotions expressed in the work, and applies to acts that are considered not to affect the opportunities to collect revenues from the work, and not to harm the interests of the copyright holder protected by the Copyright Act.

That Article also provides that the use is limited to the extent considered necessary, and it does not apply to cases where the interests of the copyright holder are unduly harmed. [...]

Re: Japan’s government will not enforce copyrights on data used in AI training

#162

Earlier quoted context omitted.

>compress an artist's painting into a model That's not how image models work.

Painting features => back propagation => weights. Yes it is.

No, it's not. It's quite easy to learn this stuff, about as easy as making shit up on a message forum, so there's no reason not to go learn what you're talking about.

Re: Japan’s government will not enforce copyrights on data used in AI training

#163
post #159
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

Pretty good argument but it has one fatal flaw. People can memorize the Declaration of Independence too. Or Harry Potter. If people mostly recite HP from memory but apply enough creative changes, it's not copyright infringement. So proving a system can memorize and recite proves nothing.

How does this make sense? Memorizing and then reciting copyrighted works is still infringement in a lot of commercial contexts.

Re: Japan’s government will not enforce copyrights on data used in AI training

#164

Earlier quoted context omitted.

Painting features => back propagation => weights. Yes it is.

No, it's not. It's quite easy to learn this stuff, about as easy as making shit up on a message forum, so there's no reason not to go learn what you're talking about.

Why don't you counter knowledge with knowledge instead of profanity then? We're all waiting.

Re: Japan’s government will not enforce copyrights on data used in AI training

#165
post #34

Earlier quoted context omitted.

I've been amused by the WH40k videos narrated by David Attenborough. There was some discussion about this a few years ago regarding Lyrebird - https://news.ycombinator.com/item?id=14182580 In particular, celebrities have an additional right - Right of Publicity. https://www.law.cornell.edu/wex/publicity > In the United States, the right of publicity is largely protected by state common or statutory law. Only about ha…

> I've been amused by the WH40k videos narrated by David Attenborough. Have you seen the Thomas the Tank Engine videos narrated by Ringo Starr and George Carlin? …

Heh Heh Heh

https://www.youtube.com/watch?v=2a_gW1KvuFk

Re: Japan’s government will not enforce copyrights on data used in AI training

#166
post #159
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

Pretty good argument but it has one fatal flaw. People can memorize the Declaration of Independence too. Or Harry Potter. If people mostly recite HP from memory but apply enough creative changes, it's not copyright infringement. So proving a system can memorize and recite proves nothing.

These aren’t people. Just because we can find commonalities in learning and memorization does not mean we can ignore everything else that differs.

Re: Japan’s government will not enforce copyrights on data used in AI training

#167
post #163
post #159

Earlier quoted context omitted.

Pretty good argument but it has one fatal flaw. People can memorize the Declaration of Independence too. Or Harry Potter. If people mostly recite HP from memory but apply enough creative changes, it's not copyright infringement. So proving a system can memorize and recite proves nothing.

How does this make sense? Memorizing and then reciting copyrighted works is still infringement in a lot of commercial contexts.

reciting is violation of copyright

creatively transform and apply for some tasks maybe not violation

Re: Japan’s government will not enforce copyrights on data used in AI training

#168

Earlier quoted context omitted.

> I've been amused by the WH40k videos narrated by David Attenborough. Have you seen the Thomas the Tank Engine videos narrated by Ringo Starr and George Carlin? …

Heh Heh Heh https://www.youtube.com/watch?v=2a_gW1KvuFk

… Now, that’s the one where they remix his narration on the show with his standup routines. It’s still mind blowing that they actually gave him the part to begin with. Same for Ringo.

Re: Japan’s government will not enforce copyrights on data used in AI training

#169

Earlier quoted context omitted.

Heh Heh Heh https://www.youtube.com/watch?v=2a_gW1KvuFk

… Now, that’s the one where they remix his narration on the show with his standup routines. It’s still mind blowing that they actually gave him the part to begin with. Same for Ringo.

The Warhammer 40k stuff with David Attenborough seems pretty interesting too.

eg: https://www.youtube.com/watch?v=x_XAhAfcTWs

Re: Japan’s government will not enforce copyrights on data used in AI training

#170
post #159
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

Pretty good argument but it has one fatal flaw. People can memorize the Declaration of Independence too. Or Harry Potter. If people mostly recite HP from memory but apply enough creative changes, it's not copyright infringement. So proving a system can memorize and recite proves nothing.

"copying" != "copyright infringement": I'm just saying that the LLMs are copying, and I'm not getting into the legal/societal question of whether we want that to be illegal or not.

We as a society have determined that certain sorts of non-consensual copying are allowed: "fair use" broadly, and maybe you can consider "mental copying" in this category. Maybe we'll add LLM training to the list? It's not like copyright rules are a law of nature: we created them to try to produce the society that we want, and this is an ongoing process.

Again, I think there are fascinating questions 1) does LLM training violate existing copyright law + case law or does it maybe fall under a fair use exemption, and 2) is that what we want. But I think "do LLMs make copies" is dull and trivial and I don't know why it comes up.

Post reply on HN