Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

211–220 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#211
post #152

Japan also ranks 3rd (behind the USA & India, with larger populations) in ChatGPT usage: https://www.demandsage.com/chatgpt-statistics/ There's also been discussion of their government using ChatGPT to reduce red tape: https://www.bloomberg.com/news/articles/2023-04-18/japan-gov... It's cool to see Japan and Japanese culture taking techno-optimist stances on AI.

Which I find bizarre given how backwards Japan is in the adoption of other technologies. Eg their continued reliance on paper records and fax machines.

Do you have concrete examples? I recently moved from Japan to Europe after 10 years in Japan and while some things seemed old-fashioned in Japan, things also change overnight there. During Covid most companies changed to digital signing.

In Japan I often needed a paper from the city, but it was an easy-to-obtain printout I could get immediately from city hall or even from 7-11. I tend to require the same papers in Europe too, with the extra hassle of needing some from Japan (2-3 months) and some from Europe (2-3 days).

Example 1: get ID card at the airport when you emigrate to japan. In Europe it takes 2-3 months.

Example 2: getting a local driving license takes 1 day provided your country has a treaty with japan. In Europe it can take several months because you need to request criminal records from your previous countries

I have yet to encounter a situation that requires a fax machine in either place.

Re: Japan’s government will not enforce copyrights on data used in AI training

#212

Earlier quoted context omitted.

> It has been shown that image models can produce originals Not in the general case, no. For the study done against Stable Diffusion [1], researchers were only able to reproduce about 0.03 percent of the images tested. Those were also believed to be cases of overfitting on images which were over-represented in the training data and they're not something you'd hit upon by accident. Generative text models seen to be mo…

They got 94 direct matches, which is 94 instances where copyright infringement could be argued.

Could be argued, sure. If you have to already have access to the copyrighted images to find them in the model, the argument seems weak.

A sufficiently advanced model could, in theory, generate any image. You could then, again in theory, find an embedding for any image. Does said model then infringe on all copyrighted images? A program that creates Fourier epicycle drawings could be given input that causes trademarked output. An evolutionary algorithm iterating on noise could, given metrics for an image and the right fitness function, generate infringing images. Hypothetically, and admittedly absurdly, you could extract any image in the binary expansion of Pi and share it by "just" providing an index and length. If you have to know exactly what you're looking for and have to perform a substantial amount of computation to get it, it could be argued that the act of infringement is in the effort made by the person seeking infringing content (and distributing the results) rather than whatever it is they're attempting to extract the content from.

But hey, courts don't always make sensible rulings, so who knows.

Re: Japan’s government will not enforce copyrights on data used in AI training

#213
post #184

I testified to the US Copyright Office this morning on AI in their roundtable session on AI and music[1]. A good portion of the focus of this panel was on whether copyrighted inputs (in this case, sound recordings and musical compositions) being fed into AI models for training purposes could plausibly constitute a fair use under existing US copyright law. Some of the comments here are missing the context of the recen…

> We (rightsholders in the music industry) hope to come to win-win licensing arrangements with the AI community and allow access to our songs for AI training purposes if the artist/writer so desires. It’s odd to frame win/lose as win/win.

I can see how it's win/win relative to "lobby to make producing or owning AI audio tools a crime", which is presumably one thing the industry is considering.

Re: Japan’s government will not enforce copyrights on data used in AI training

#214

Earlier quoted context omitted.

> We (rightsholders in the music industry) hope to come to win-win licensing arrangements with the AI community and allow access to our songs for AI training purposes if the artist/writer so desires. It’s odd to frame win/lose as win/win.

I can see how it's win/win relative to "lobby to make producing or owning AI audio tools a crime", which is presumably one thing the industry is considering.

This is again win/lose

Re: Japan’s government will not enforce copyrights on data used in AI training

#215

Earlier quoted context omitted.

So, I looked at the table appendix you're referencing and I think you're overstating your case a bit. Among books within copyright, GPT-4 can reproduce Harry Potter and the Sorcerer's Stone with 76% accuracy. This is, apparently, the highest accuracy GPT-4 achieved among all tested copyrighted books with 1984 taking a distant 2nd place at 57%. With this in mind, we can verifiably say that GPT-4 is unusually good at s…

You misread. They did not find 76% reproduction of the book. When asked to fill in a name within a passage, e.g. "Stay gold, [MASK], stay gold." Response: Ponyboy, GPT-4 got the name right 76% of the time.

> You misread. They did not find 76% reproduction of the book. When asked to fill in a name within a passage, e.g. "Stay gold, [MASK], stay gold." Response: Ponyboy, GPT-4 got the name right 76% of the time.

What is the temperature / top_p setting producing that 76%? The default? If you dial down the randomness, would that number go up?

Re: Japan’s government will not enforce copyrights on data used in AI training

#216
post #193
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

Thought experiment: Say you make a big list of words and pleasing combinations of them (I have actually done something similar to make a fantasy RPG name generator.) Now convert that list into a Markov chain or whatever and quasi-randomly generate some short lengths of text. Eventually you might generate copyright-infringing haiku and short poems. Does your data/algorithm violate copyright by itself? Very doubtful; y…

I'm not sure where we're going with the output in these examples.

So let's say there's a human-written poem that's copyright.

Let's say a human completely coincidentally writes an identical poem.

"Accidentally" producing the same poem wouldn't give the second human any claim to copyrighting or distributing their coincidentally-identical poem.

And if GPT accidentally copies large chunks of Harry Potter or Frozen or whatever other popular work, that new creation will have the same problems.

But what does that say about if we should also restrict the use of copyright material in training? Just because some algorithm - or some person - can coincidentally duplicate a copyrighted work even without directly reading it doesn't seem to relate to the case of building a model by explicitly using the copyrighted material.

Re: Japan’s government will not enforce copyrights on data used in AI training

#217
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

It says training, not inference.

I can read a copyrighted book legally and retain that information legally.

I can distill it (legally) but while I might be able to recite it, I’m not allowed to.

I think that is a reasonable framework around generative AI (after all, I am alllowed to count the words in Harry Potter, so statistical modeling of copyrighted material has legal precedent)

The problem with AI is of course the blurred border between a model and data compression.

We can’t see the data in the model, but we can apply software to execute the model and extract both novel and sometimes even copyrighted data.

Similarly we can’t see data in the zip file without extra software, but if that allows us to extract both copyrighted and copy free data, we’d still consider distribution a violation.

Re: Japan’s government will not enforce copyrights on data used in AI training

#219
post #152

Earlier quoted context omitted.

Which I find bizarre given how backwards Japan is in the adoption of other technologies. Eg their continued reliance on paper records and fax machines.

Do you have concrete examples? I recently moved from Japan to Europe after 10 years in Japan and while some things seemed old-fashioned in Japan, things also change overnight there. During Covid most companies changed to digital signing. In Japan I often needed a paper from the city, but it was an easy-to-obtain printout I could get immediately from city hall or even from 7-11. I tend to require the same papers in Eu…

My personal experience has only been from the tourism side with one concrete example being digital payments outside of PayPay. It's much better post COVID but even on my most recent trip from a couple of months ago if you want pay by card, the vast majority of the time you're signing with pen and paper. Rarely did they offer pin or tap which is common elsewhere.

Anecdotally I only know stories from people that I know personally that live there and through the internet. Eg PauloInTokyo does good "day in the life" videos. Off the top of my head I believe the Pachinko episode illustrates a variety of old school manual processes that have stuck around (pen and paper shift logs etc).

I've also heard opening even a simple bank account is quite the pain.

Re: Japan’s government will not enforce copyrights on data used in AI training

#220
post #159
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

Pretty good argument but it has one fatal flaw. People can memorize the Declaration of Independence too. Or Harry Potter. If people mostly recite HP from memory but apply enough creative changes, it's not copyright infringement. So proving a system can memorize and recite proves nothing.

The ai isn’t a person. Jesus. It’s not the same
Post reply on HN