Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

191–200 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#191
post #184

I testified to the US Copyright Office this morning on AI in their roundtable session on AI and music[1]. A good portion of the focus of this panel was on whether copyrighted inputs (in this case, sound recordings and musical compositions) being fed into AI models for training purposes could plausibly constitute a fair use under existing US copyright law. Some of the comments here are missing the context of the recen…

How do you feel about human musicians learning from copyrighted works? Technical limitations aside, is that something you'd like to monetize?

Re: Japan’s government will not enforce copyrights on data used in AI training

#192

Earlier quoted context omitted.

Notably none of those things, if they did apply to the art student, are copyright violations. The speed, scale, versatility, and ownership of a machine learning model has no bearing on its ability to violate copyright.

Good reasons not to assign copyright to their output, at least not without some caveats.

I don't have strong feelings about that either way, but that's not what this post is about.

Re: Japan’s government will not enforce copyrights on data used in AI training

#193
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

Thought experiment: Say you make a big list of words and pleasing combinations of them (I have actually done something similar to make a fantasy RPG name generator.) Now convert that list into a Markov chain or whatever and quasi-randomly generate some short lengths of text. Eventually you might generate copyright-infringing haiku and short poems. Does your data/algorithm violate copyright by itself? Very doubtful; you wrote it all yourself. Only publishing the output violates copyright. (See also: http://allthemusic.info/)

So if that's legal, how about if, instead of entering the data manually, you write an algorithm to scan poetry and collect statistics about the words in it. Should the legal distinction be any different since all you did was automate the manual process above?

Or what if you used a big list of the titles of poetry, which isn't even copyrightable information by itself? You may still succeed in extracting the aesthetic intent of the authors, and a statistical model can plausibly use that to generate copyright-infringing work.

Remember, we're not talking about generating novels or paintings here, just 20 words or so (whatever the bare minimum copyrightable amount is) in trillions of generated permutations.

You can see where I'm going with this. If those examples are legal, is there a cut-off for more complex statistical systems? Good luck figuring that out in a court of law.

Re: Japan’s government will not enforce copyrights on data used in AI training

#194

Earlier quoted context omitted.

Sure, but the key part there is "some of". They're necessarily able to produce verbatim copies only of the most duplicated, most repeated, most cited works -- and it's precisely due to their popularity that they're the only things worth including verbatim. I'm not going to opine on what the legality of that should be, but it's essentially the material considered most "quotable" in different contexts. I'm quite sure t…

> I'm quite sure the entirety of Harry Potter isn't included, but I'm also sure that some of the most popular paragraphs probably are. It's analagous to the kind of stuff people memorize. No, you are wrong about this. There are good reasons to believe the model memorized the entirety of Harry Potter, as well as Fifty Shades of Grey, inclusive of unremarkable paragraphs, the kind of stuff people will never memorize. B…

So, I looked at the table appendix you're referencing and I think you're overstating your case a bit.

Among books within copyright, GPT-4 can reproduce Harry Potter and the Sorcerer's Stone with 76% accuracy. This is, apparently, the highest accuracy GPT-4 achieved among all tested copyrighted books with 1984 taking a distant 2nd place at 57%.

With this in mind, we can verifiably say that GPT-4 is unusually good at specifically reproducing the first Harry Potter book. An unscrupulous book thief may very well be able to steal the first entry in the series... assuming that they're able to get past one quarter of the book being an AI hallucination.

Re: Japan’s government will not enforce copyrights on data used in AI training

#195
post #184

I testified to the US Copyright Office this morning on AI in their roundtable session on AI and music[1]. A good portion of the focus of this panel was on whether copyrighted inputs (in this case, sound recordings and musical compositions) being fed into AI models for training purposes could plausibly constitute a fair use under existing US copyright law. Some of the comments here are missing the context of the recen…

[flagged]

Re: Japan’s government will not enforce copyrights on data used in AI training

#196

So if you train an audio model on say, Eminem's voice, then write some songs and have it perform them...Would this output be legal to publish?

I hope smarter minds than myself will somehow figure out a way to square this circle and get heavy restrictions on commercial usage without killing the technology outright.

As far as AI rap goes: Notorious BIG is (by far) my favorite rapper of all time, yet he died in his early 20s with barely any unreleased songs in the vault. For me, as an aging superfan, AI deepfakes have me feeling 15 again - like Biggie is still alive somewhere and dropping covers of other classic rap tracks on YouTube.

I wanted to link an example of an AI Biggie cover of a Nas classic, but the full thing seems to have vanished off YT after getting some press a few weeks back. A Shorts snippet is all I could find quickly.[1]

And if you’re not a fan of old rap, these Michael Jackson[2] and Freddy Mercury[3] “AI cover songs” are absolutely wild.

I just don’t see how this sort of stuff holds zero cultural value. Yes, it’s fake and a cover song to boot - but if the music elicits an emotional response that differs from both “original” artists of the deepfake, isn’t that art?

1. https://youtube.com/shorts/l81JqY2uhEo

2. https://youtu.be/370dSdNRYG4

3. https://youtu.be/XiZtIARF0iM

Re: Japan’s government will not enforce copyrights on data used in AI training

#197
post #193
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

Thought experiment: Say you make a big list of words and pleasing combinations of them (I have actually done something similar to make a fantasy RPG name generator.) Now convert that list into a Markov chain or whatever and quasi-randomly generate some short lengths of text. Eventually you might generate copyright-infringing haiku and short poems. Does your data/algorithm violate copyright by itself? Very doubtful; y…

> Remember, we're not talking about generating novels or paintings here, just 20 words or so (whatever the bare minimum copyrightable amount is)

From https://fairuse.stanford.edu/2003/09/09/copyright_protection...:

Copyright laws disfavor protection for short phrases. Such claims are viewed with suspicion by the Copyright Office, whose circulars state that, “… slogans, and other short phrases or expressions cannot be copyrighted.” [1] These rules are premised on two tenets of copyright law. First, copyright will not protect an idea. Phrases conveying an idea are typically expressed in a limited number of ways and, therefore, are not subject to copyright protection. Second, phrases are considered as common idioms of the English language and are therefore free to all. Granting a monopoly would eventually “checkmate the public” [2] and the purpose of a copyright clause to encourage creativity-would be defeated.

Re: Japan’s government will not enforce copyrights on data used in AI training

#198
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

Typically things like this are covered under fair use if you're dealing with a human.

Re: Japan’s government will not enforce copyrights on data used in AI training

#199

Earlier quoted context omitted.

> I'm quite sure the entirety of Harry Potter isn't included, but I'm also sure that some of the most popular paragraphs probably are. It's analagous to the kind of stuff people memorize. No, you are wrong about this. There are good reasons to believe the model memorized the entirety of Harry Potter, as well as Fifty Shades of Grey, inclusive of unremarkable paragraphs, the kind of stuff people will never memorize. B…

So, I looked at the table appendix you're referencing and I think you're overstating your case a bit. Among books within copyright, GPT-4 can reproduce Harry Potter and the Sorcerer's Stone with 76% accuracy. This is, apparently, the highest accuracy GPT-4 achieved among all tested copyrighted books with 1984 taking a distant 2nd place at 57%. With this in mind, we can verifiably say that GPT-4 is unusually good at s…

You misread. They did not find 76% reproduction of the book. When asked to fill in a name within a passage, e.g. "Stay gold, [MASK], stay gold." Response: Ponyboy, GPT-4 got the name right 76% of the time.

Re: Japan’s government will not enforce copyrights on data used in AI training

#200
post #193

Earlier quoted context omitted.

Thought experiment: Say you make a big list of words and pleasing combinations of them (I have actually done something similar to make a fantasy RPG name generator.) Now convert that list into a Markov chain or whatever and quasi-randomly generate some short lengths of text. Eventually you might generate copyright-infringing haiku and short poems. Does your data/algorithm violate copyright by itself? Very doubtful; y…

> Remember, we're not talking about generating novels or paintings here, just 20 words or so (whatever the bare minimum copyrightable amount is) From https://fairuse.stanford.edu/2003/09/09/copyright_protection... : Copyright laws disfavor protection for short phrases. Such claims are viewed with suspicion by the Copyright Office, whose circulars state that, “… slogans, and other short phrases or expressions cannot b…

You could still plausibly generate (a significant portion of), let's say, "Fire And Ice" by Robert Frost, which is only 50 words.

See also: https://blogs.harvard.edu/ethicalesq/haiku-and-the-fair-use-...

Post reply on HN