Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

221–230 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#221
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

> Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?)

This would make a fun short story - “ChatGPT, author of the Quixote”

Re: Japan’s government will not enforce copyrights on data used in AI training

#222
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

You don't need to even read their law to know that they are speaking only of training and not of output. Otherwise, they would have just suddenly created the world's most obvious loophole. Create an 'LLM' that "trains" on some input and then categorically outputs each file, be it a movie, song, book, or whatever. You've now legalized copyright infringement (and distribution) of everything.

So their law is going to essentially come down to you can train your LLM on whatever you want, but can also be held liable for any infringing outputs.

Re: Japan’s government will not enforce copyrights on data used in AI training

#223
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

The things that an LLM is likely to contain a complete verbatim copy of are things that are a) short b) widely repeated to the point that they're embedded into our culture - and by that token those things are almost certainly not copyrightable.

Re: Japan’s government will not enforce copyrights on data used in AI training

#224
post #140
post #89

Earlier quoted context omitted.

I strongly agree with this. There's a distinction between "learning from" and "copying". "Learning from" is a transformative process that distills from the observation. This distillation can be as simple as indexing for a search engine, or as complex as a deep neural network. Simply because a neural network can create something that is a copyright violation doesn't mean the training process itself it. A human can see…

> A human can see a advertisement for a Marvel movie and then reproduce the Marvel logo. Redistributing (and possibly actually doing that reproduction) that logo is a copyright violation, but the learning process isn't. I don't think that's correct. That might be trademark infringement, if the logo is a registered trademark, but "seeing something and then drawing it" is in general not copyright infringement.

Drawing a copy of a copyrighted picture from memory, and then distributing that copy, would certainly normally be copyright infringement. (A logo may not be enough of a creative work to be copyrightable, but I assume that's not what you're getting at).

Re: Japan’s government will not enforce copyrights on data used in AI training

#225
post #223
post #155

Not trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see ho…

The things that an LLM is likely to contain a complete verbatim copy of are things that are a) short b) widely repeated to the point that they're embedded into our culture - and by that token those things are almost certainly not copyrightable.

Is a bar in a song not "short"?

Try putting one of those in your book and not getting sued for copyright.

Re: Japan’s government will not enforce copyrights on data used in AI training

#226
post #140
post #89

Earlier quoted context omitted.

I strongly agree with this. There's a distinction between "learning from" and "copying". "Learning from" is a transformative process that distills from the observation. This distillation can be as simple as indexing for a search engine, or as complex as a deep neural network. Simply because a neural network can create something that is a copyright violation doesn't mean the training process itself it. A human can see…

> A human can see a advertisement for a Marvel movie and then reproduce the Marvel logo. Redistributing (and possibly actually doing that reproduction) that logo is a copyright violation, but the learning process isn't. I don't think that's correct. That might be trademark infringement, if the logo is a registered trademark, but "seeing something and then drawing it" is in general not copyright infringement.

> "seeing something and then drawing it" is in general not copyright infringement.

It's seeing something, drawing it, and then distributing that drawing which is infringement. Bonus points for the distribution being a sale.

Re: Japan’s government will not enforce copyrights on data used in AI training

#227
post #163

Earlier quoted context omitted.

How does this make sense? Memorizing and then reciting copyrighted works is still infringement in a lot of commercial contexts.

The reciting part is illegal, but as long as it is trained not to recite things in full (or to whatever limit the law determines), then it should be fine.

Try publishing Harry Potter but changing all the proper nouns and use synonyms for all the adjectives.

It's gonna be copyright infringement.

You can even cut a few scenes and make up a few scenes entirely, too. You're still getting busted.

Re: Japan’s government will not enforce copyrights on data used in AI training

#228
post #51
post #22

Earlier quoted context omitted.

There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.

I've been wondering why this argument's not been sitting with me, and I think it's for the same reason that the courts have ruled that the FBI needed a warrant to put a tracker on someone's car, as opposed to following someone - the scale of action enabled is the differentiator. A student learning from other artists is still limited in their output to human-scale - they must physically create the new thing. An AI mod…

I do agree with you, but honestly I don't even think that's the biggest problem with these arguments.

I'm just sitting here wondering why it is even relevant whether the "AI" is "copying", "learning", "thinking", or whatever, why is any of that important? Does AI have human rights? Well, perhaps in a couple hundred years, if humanity manages not to self-extinguish by then.

It's not like you can sue AI if you think it plagiarized your work, no. Obviously not, so why the hell are we discussing that? "AI" is just a piece of software, a tool, it doesn't matter what it's doing, what matters is what the user is doing, the fact of the matter is that these multi-billionaire corporations are taking everyone's honest work, putting it into a computer, and selling the output. They didn't do any "learning", they just used your data and made money out of it, it isn't a stretch to say they simply sold your work.

EDIT: Perhaps one day the day AI will have human rights, make its own money, and pay bills. That will be the day any of this nonsensical discussion will be anything but useless.

Re: Japan’s government will not enforce copyrights on data used in AI training

#229
post #138

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

> The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. Lossy or not, the training data provides value . If all the various someones had not spent time making all the stuff that ends up as training data, then the model it trains would not exist. If you are going to use someone else's work in order to make something that you are going to…

> If you are going to use someone else's work in order to make something that you are going to profit off of, I believe that original author should be compensated. And should also be able to decide they don't want their work used in that way.

> Note that I'm not talking about what existing copyright law says; I'm talking about how I believe we should be regulating this new facet of the industry.

Is it really new? Humans have always learnt by studying what's out there already. Our whole culture is built on what's been done and published before (and how could it be otherwise?). Without Bach there would be no Mozart, and then down the line their influence permeates everything you hear today.

If anything I'd like to make it easier to reuse parts of our shared culture, and limit the ability of organisations to control how things that they've published are reused. You can make private works if you want to keep control of them, but at some point the public deserves to share and rework the things that have been pushed into the public consciousness.

Re: Japan’s government will not enforce copyrights on data used in AI training

#230
post #80
post #22

Earlier quoted context omitted.

There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.

Conflating training a model with human learning is wrong. When training a model you are deriving a function that takes some input and produces an output. The issue with copyright and licensing here is that a copy is made and reproduced numerous times when training. The model is not walking around a museum where it is an authorized viewing. It is not a being learning a skill. It is a function. The further issue is tha…

> When training a model you are deriving a function that takes some input and produces an output. The issue with copyright and licensing here is that a copy is made and reproduced numerous times when training.

How's that any different from what happens inside a human's brain when learning?

> The model is not walking around a museum where it is an authorized viewing.

The training data could well be from an online museum. And the idea that viewing something public has to be "authorized" is very insidious.

> The further issue is that it may output material that competes with the original.

So might a human student.

Post reply on HN