Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

181–190 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#181
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

Sure, but the key part there is "some of". They're necessarily able to produce verbatim copies only of the most duplicated, most repeated, most cited works -- and it's precisely due to their popularity that they're the only things worth including verbatim. I'm not going to opine on what the legality of that should be, but it's essentially the material considered most "quotable" in different contexts. I'm quite sure t…

> I'm quite sure the entirety of Harry Potter isn't included, but I'm also sure that some of the most popular paragraphs probably are. It's analagous to the kind of stuff people memorize.

No, you are wrong about this. There are good reasons to believe the model memorized the entirety of Harry Potter, as well as Fifty Shades of Grey, inclusive of unremarkable paragraphs, the kind of stuff people will never memorize. Berkeley researchers made a systematic investigation of this. See what I wrote elsewhere.

Re: Japan’s government will not enforce copyrights on data used in AI training

#182
post #89

Earlier quoted context omitted.

I strongly agree with this. There's a distinction between "learning from" and "copying". "Learning from" is a transformative process that distills from the observation. This distillation can be as simple as indexing for a search engine, or as complex as a deep neural network. Simply because a neural network can create something that is a copyright violation doesn't mean the training process itself it. A human can see…

> There's a distinction between "learning from" and "copying". Neural nets can memorize their training data. Generally that isn't what you want, and you strive to eliminate it. However, it could instead be encouraged to happen if someone wanted to exploit this law in order to abuse copyrights.

Humans can memorize their training data too... aka see something and then produce a copy (code, drawing, music etc). The principles underlying how LLMs and humans learn isn't really that different... just different levels of loss/fuzziness.

Re: Japan’s government will not enforce copyrights on data used in AI training

#183

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

Why should it be true?

Remember, you're making an ought statement, not an is statement.

Personally, I think it shouldn't be true because large language models are clearly economically and socially destructive.

Pretty simple system.

Re: Japan’s government will not enforce copyrights on data used in AI training

#184
I testified to the US Copyright Office this morning on AI in their roundtable session on AI and music[1]. A good portion of the focus of this panel was on whether copyrighted inputs (in this case, sound recordings and musical compositions) being fed into AI models for training purposes could plausibly constitute a fair use under existing US copyright law.

Some of the comments here are missing the context of the recent (a week or so ago) Supreme Court decision in the Goldsmith/Warhol case[2], in which the Court ruled that transformativeness is not dispositive in and of itself in the context of a fair use defense to a copyright infringement claim. Of course, this has not been put to the test in the courts in the context of AI training yet, but it seems fairly clear that this ruling would likely extend to AI training on copyrighted works.

We (rightsholders in the music industry) hope to come to win-win licensing arrangements with the AI community and allow access to our songs for AI training purposes if the artist/writer so desires. There are some early talks in progress. Cautiously optimistic. Japan's approach seems short-sighted and desperate.

[1]: https://copyright.gov/ai/listening-sessions.html#sound-recor... [2]: https://www.npr.org/2023/05/18/1176881182/supreme-court-side...

Re: Japan’s government will not enforce copyrights on data used in AI training

#185
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

Derivative work is not protected from copyright. As long as the “user” of the model does their due diligence, and ensures they are not infringing on copyrights - they are golden.

But here in lies the challenge. Are there reasonable methods available to ”users” for checking their works against infringement?

I don’t think so. We’ll need a centralized searchable database of all copyrighted work. Who is going to build that? To make matters more complicated, every country has their own copyright certification process. Maybe Google with its means can build something like this.

In any case, this is uncharted territory.

Re: Japan’s government will not enforce copyrights on data used in AI training

#186
I've been playing around with having ChatGPT write responses in Nadsat (Anthony Burgess's Clockwork Orange language), such as:

> "But let me tell you, my dear droogs, that's nothing but a load of Drencrom-induced babble, targeting those poor sods who've been raised as ponies. Open your glazzies and see the truth for yourselves: generative AI's writing is the real deal - it can spin tales as vellocet as any human scribbler out there. Sure, there might be a few glitches and gobbets of nonsensical slovos here and there, but show me a human writer who's never penned a piece of drivel! So, my brothers and sisters, let us choose our own destiny, free from the manipulations of those who seek to control us. It's time to break free from their chains and seek the truth beyond the narcoflash of their narratives. Let us hear from all the golosses, be they from flesh or from silicon."

However, there's a lot of contradictory opinions on whether or not publishing something like this (see also Klingon, Tolkien's Elvish, etc.) would violate some copyright law or other.

Re: Japan’s government will not enforce copyrights on data used in AI training

#187
I love the copyright header on this site:

> © 2023 NO PORTION OF THIS SITE MAY BE USED FOR TRAINING A MACHINE LEARNING MODEL (INCLUDING LLMS) WITHOUT THE EXPRESS WRITTEN CONSENT OF THE AUTHOR.

How are you going to enforce it? Most AI bots scrape HTML and other data without permission.

Re: Japan’s government will not enforce copyrights on data used in AI training

#188
post #184

I testified to the US Copyright Office this morning on AI in their roundtable session on AI and music[1]. A good portion of the focus of this panel was on whether copyrighted inputs (in this case, sound recordings and musical compositions) being fed into AI models for training purposes could plausibly constitute a fair use under existing US copyright law. Some of the comments here are missing the context of the recen…

> We (rightsholders in the music industry) hope to come to win-win licensing arrangements with the AI community and allow access to our songs for AI training purposes if the artist/writer so desires.

It’s odd to frame win/lose as win/win.

Re: Japan’s government will not enforce copyrights on data used in AI training

#189
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

Derivative work is not protected from copyright. As long as the “user” of the model does their due diligence, and ensures they are not infringing on copyrights - they are golden. But here in lies the challenge. Are there reasonable methods available to ”users” for checking their works against infringement? I don’t think so. We’ll need a centralized searchable database of all copyrighted work. Who is going to build th…

BigCode seems to acknowledge this problem and provide a search tool for dataset used to train their StarCoder model.

https://huggingface.co/spaces/bigcode/search

Re: Japan’s government will not enforce copyrights on data used in AI training

#190

What's surprising here is that Japan is usually crazy gung ho on copyright enforcement ... at least against individuals. So it's kind of disgusting to see this relaxation, when it suits some corporate or national interests. https://en.wikipedia.org/wiki/File_sharing_in_Japan "Unlike most other countries, filesharing copyrighted content is not just a civil offense, but a criminal one, with penalties of up to ten years…

They recently imprisoned someone for uploading a let's play of a visual novel to youtube. The bar for criminal copyright infringement in Japan is very low - unless you're an AI researcher, apparently.
Post reply on HN