Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

411–420 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#411

Earlier quoted context omitted.

> If you have to already have access to the copyrighted images to find them in the model, the argument seems weak. That makes no sense. The copyright holder has access to their own inventions, of course. That's the standard in any copyright claim. > A sufficiently advanced model could, in theory, generate any image. You could then, again in theory, find an embedding for any image. Does said model then infringe on all…

> That makes no sense. The copyright holder has access to their own inventions, of course. No, not talking about the copyright holder, I'm talking about the hypothetical individual(s) creating infringing copies. If those people need to already have a copy of the image to extract a copy of the image from the generative image model, then I'm saying the argument that the model itself is infringing seems weak. Or it's at…

> If those people need to already have a copy of the image to extract a copy of the image from the generative image model, then I'm saying the argument that the model itself is infringing seems weak.

AFAIK, they don't need that to extract it.

> Instead, you already have to have a copy of an image to find an embedding.

You literally always need a copy or a suitable imprecise hash of the original to test for infringement. How else would you know which copyright was infringed? But it's not a matter of a 1-1 match (see below).

> However, it's a case-by-case thing.

Of course, it's a case-by-case thing, as the OP also indicated. It's easy to show mathematically that these models cannot compress well enough to contain all training images. The question is how much they infringe on some of them.

Bear in mind that it is not necessary at all to create a perfect copy of an image the infringe copyright. As I've stated above, even a lousy and mostly incorrect rendition of a pop song in a street cafe may infringe copyright. The makers behind the song "Blurred Lines" lost a lawsuit because the cowbell rhythm in the background was similar to that of another song. That and the "feeling" was similar.

The same is true for images. What counts are criteria like artistic originality, intent, subjective similarities, experts laying out similarities in style, and so on.

I mean, don't get me wrong, I understand perfectly well what you're trying to argue for. All I'm saying is that it doesn't match the reality of how the law deals with copyright.

Re: Japan’s government will not enforce copyrights on data used in AI training

#412
post #66
post #22

Earlier quoted context omitted.

There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.

One day we'll have a thread about AI where someone doesn't use the "machines deserve the same rights as people" non-argument. But this isn't that thread.

Actually, the people who say things like this are arguing for the degradation of human rights, because there's significant overlap between this and "humans aren't special" and "AI is our successor species". It's nihilism all the way down but they're forcing the rest of the world on their little suicide charge and expecting everyone else to be just as enthusiastic about it as they are.

Re: Japan’s government will not enforce copyrights on data used in AI training

#413
post #331
post #126

Earlier quoted context omitted.

Privacy, where there's a problem if some original data can be inferred/made out (even using a white box attack against the model), is a higher bar than whether an image generator avoids copyright-infringing output under non-adversarial usage. Additionally, compared to image data, text is more prone to exact matches due to lower dimensionality and usually training with less data per parameter. While it's still a topic…

The privacy issue isn’t just about the data being available as people’s names, addresses, and phone numbers are generally available. The issue is if they show up as part of some meme chat and then you as the LLM creator get sued because people start harassing them. In terms of copyright infringement the bar is quite low, and copying is a basic part of how these algorithms work. This may or may not be an issue for you…

> The issue is if they show up as part of some meme chat and then you as the LLM creator get sued because people start harassing them.

This seems a more obscure concern than extraction of data.

> copying is a basic part of how these algorithms work

Do you mean during training/gradient descent, or reverse diffusion?

Re: Japan’s government will not enforce copyrights on data used in AI training

#414
So let's say that just hypothetically, I train an image generating model with every available work drawn or designed by Akira Toriyama, then I use this model to make a Journey to the West comedy manga.

The government won't arrest me for that, but can private corporations still sue me for using their work to train the model?

Re: Japan’s government will not enforce copyrights on data used in AI training

#415

Earlier quoted context omitted.

Anything you create comes from what you've seen, from what you've experienced. That's a dead-end to start having to pay each and every one of the original sources of the components of your mind each time they are used to create!

So I understand you are arguing that derived works should not be subject to royalties, i.e. I should be able to produce a film based entirely on the work of a living author without restriction or the need to pay any royalties.

No, you're over-generalizing what I'm saying (a.k.a. straw man fallacy).

If you're commercializing e.g. goodies that use the exact same character designed by some artist, so that anybody can tell you it's the same, AND if the traits of this character are actually original (not just so that anybody would come up with it independently), AND if the people buying your goodies are all thinking about the original character in the first place, AND if the original author is alive, then I'd absolutely agree that royalties need to be paid.

That's just an example. To tell you that in specific cases it's obvious that royalties are required.

But in the general case, no. Because, else, anything just is a derived work. Just think about it.

Re: Japan’s government will not enforce copyrights on data used in AI training

#416

Earlier quoted context omitted.

So I understand you are arguing that derived works should not be subject to royalties, i.e. I should be able to produce a film based entirely on the work of a living author without restriction or the need to pay any royalties.

No, you're over-generalizing what I'm saying (a.k.a. straw man fallacy). If you're commercializing e.g. goodies that use the exact same character designed by some artist, so that anybody can tell you it's the same, AND if the traits of this character are actually original (not just so that anybody would come up with it independently), AND if the people buying your goodies are all thinking about the original character…

It's even impossible to retrace all of the woven threads, ramified tendrils, that link your mind to the billions of other minds all over the world and over the ages.

The world of ideas is liquid. All is mixing, all dissolves and disappears in everything else, and is reborn new and different, again and again.

Re: Japan’s government will not enforce copyrights on data used in AI training

#417
post #317

Earlier quoted context omitted.

The same way you do with any other content you generate in other ways.

Normally when I generate content it’s from my brain and I can tell the difference between copying memorized content, re-expressing memorized content, and generating something original. How do I know what the LLM is doing?

Are you sure? If you look at plagiarism in music, you'll find a number of cases where the defendant makes a compelling point about not remembering or consciously knowing they heard the original song before. For legal purposes, it is not the point, but they feel morally wronged to be charged as guilty. The case here is that they internalized the music knowledge, but forgot about the source - so they can't make the distinction you claim anymore. Natural selection shaped our brains to store formation that seems useful, not is attribution.

LLMs are also not usually trained to remenber where the examples they were trained on came from, the sourcing information is often not even there (maybe they could, maybe they should, but they aren't). Given that and the way training works, one could argue that they're never copying, only re-expressing or combining (which I think of as a form of "generating something original"). Just memorizing and copying is overfitting, and strongly undesirable, as it's not usable outside of the exact source context. I agree it can happen, but it's a flaw in the training process. I'd also agree that any instance of exact reproductions (or of material with similarity to the original content over some high threshold) is indeed copyright infringement, punishable as such.

So, my point is, training a model on copyrighted material is legal, but letting that model output copies of copyrighted material beyond fair use (quotations, references, etc - that make sense in the context the model was queried on) is an infringement. And since the actual training data is not necessarily known, providers of model-as-a-service, such as OpenAI with GPT, should be responsible for that.

In cases where a model was made available to others, it falls on the user of the model. If the training data is available, they should check answers against it (there's a whole discussion on how training data should be published to support this) to avoid the risk;if the training data is unknown, they're taking the risk of being sued full-on, without any mitigation.

Re: Japan’s government will not enforce copyrights on data used in AI training

#418
post #91

Earlier quoted context omitted.

> There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. A few notable differences: 1. Scale: a single art student can't view millions of works in a week. 2. Duplication: a single art student's brain can't be cloned or downloaded into another art student's brain. 3. Speed: a single art student cannot draw or paint thousands of images in a…

These are all true. But they apply equally to a search engine index and that has already been found not to violate copyright and to be very useful to society.

> These are all true

That was the point.

Whether or not there is copyright infringement going on (or whether copyright law is an appropriate regulatory framework for ML models), I frequently see claims like "ML training is no different from what humans do" repeated, in spite of its incorrectness.

Re: Japan’s government will not enforce copyrights on data used in AI training

#419

Earlier quoted context omitted.

> There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. A few notable differences: 1. Scale: a single art student can't view millions of works in a week. 2. Duplication: a single art student's brain can't be cloned or downloaded into another art student's brain. 3. Speed: a single art student cannot draw or paint thousands of images in a…

Notably none of those things, if they did apply to the art student, are copyright violations. The speed, scale, versatility, and ownership of a machine learning model has no bearing on its ability to violate copyright.

The interesting thing to me is determining when something that is acceptable for humans to do at human scale is also acceptable for machines to do at industrial scale.

There are many examples. One is face recognition - clearly it is acceptable for individual humans to do this at small scale, but systematic identification of everyone on a street or in a stadium has different implications for society.

(In this case it almost doesn't matter whether the surveillance is performed by humans or machines - it's the scale and systematic nature that changes the equation.)

Re: Japan’s government will not enforce copyrights on data used in AI training

#420

Earlier quoted context omitted.

> There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. A few notable differences: 1. Scale: a single art student can't view millions of works in a week. 2. Duplication: a single art student's brain can't be cloned or downloaded into another art student's brain. 3. Speed: a single art student cannot draw or paint thousands of images in a…

If an art student was able to gain those abilities, would you then argue they couldn't legally create art?

[deleted]
Post reply on HN