Earlier quoted context omitted.
It has been shown that image models can produce originals, or at least extremely close to the originals. If the outcome is the same, what is the difference between compression/decompression vs training/generation regarding copyright?
> It has been shown that image models can produce originals Not in the general case, no. For the study done against Stable Diffusion [1], researchers were only able to reproduce about 0.03 percent of the images tested. Those were also believed to be cases of overfitting on images which were over-represented in the training data and they're not something you'd hit upon by accident. Generative text models seen to be mo…
Japan’s government will not enforce copyrights on data used in AI training
201–210 of 426 posts
Re: Japan’s government will not enforce copyrights on data used in AI training
#202> Japanese Minister of Education, Culture, Sports, Science, and Technology
Re: Japan’s government will not enforce copyrights on data used in AI training
#203The original links here, to the actual question asked and the answer by the minister: https://kiitaka.net/21312/
This is an opposition figure advocating for stronger copyright protection. He was clearly trying to make a point by asking what the current laws allows regarding the use of copyrighted materials by AI. The minister simply confirmed that no regulations are currently in place to limit that.
The whole article is a blatant lie. Neither politicians went "all in" advocating the use of copyrighted materials by AI. They just confirmed what the current laws say. The fact that they're even discussing this likely means that there will be even more regulation, not less.
Re: Japan’s government will not enforce copyrights on data used in AI training
#204Earlier quoted context omitted.
> Remember, we're not talking about generating novels or paintings here, just 20 words or so (whatever the bare minimum copyrightable amount is) From https://fairuse.stanford.edu/2003/09/09/copyright_protection... : Copyright laws disfavor protection for short phrases. Such claims are viewed with suspicion by the Copyright Office, whose circulars state that, “… slogans, and other short phrases or expressions cannot b…
You could still plausibly generate (a significant portion of), let's say, "Fire And Ice" by Robert Frost, which is only 50 words. See also: https://blogs.harvard.edu/ethicalesq/haiku-and-the-fair-use-...
I think a jury would side with my argument.
Re: Japan’s government will not enforce copyrights on data used in AI training
#205[flagged]
If you are 'learning' is nothing but taking existing drawings and storing them in an electronic format (no matter that it's a super lossy a format, for example an MP3 made from a CD is still a copy no matter how low you set the low bitrate.) then no, no you are not. Now if you learning involves training a human brain on actual techniques to replicate how someone else did something, then yes, yes you are.
Re: Japan’s government will not enforce copyrights on data used in AI training
#206I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…
I strongly agree with this. There's a distinction between "learning from" and "copying". "Learning from" is a transformative process that distills from the observation. This distillation can be as simple as indexing for a search engine, or as complex as a deep neural network. Simply because a neural network can create something that is a copyright violation doesn't mean the training process itself it. A human can see…
NNs and humans don't learn the same way - humans can fairly quickly generalise what they have learned and, most importantly, go beyond what they've learned. I haven't see that happen with neural networks or GPTs; at best, you're getting the average of what it has 'learned'. There's human learning and there's neural network 'learning' and they're a different thing.
Re: Japan’s government will not enforce copyrights on data used in AI training
#207Re: Japan’s government will not enforce copyrights on data used in AI training
#208If I compress an artist's painting into a jpeg and rehost part of it for individual t-shirt designs I am committing a crime. If I compress an artist's painting into a model & rehost what's essentially a highly flexible complete version of their painting for infinite, perpetual use of any kind ... I'm not committing a crime?
What I wonder is: Load the source code to all versions of unix, with all licenses. "write me a version of unix" Since there is no model copyright and the result was written by AI the software is now in the public domain.
No that's not how copyright works. If you have photographic memory and reproduce a work exactly you still commit copyright infringement.
Re: Japan’s government will not enforce copyrights on data used in AI training
#209I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…
> The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. Lossy or not, the training data provides value . If all the various someones had not spent time making all the stuff that ends up as training data, then the model it trains would not exist. If you are going to use someone else's work in order to make something that you are going to…
is it? If i distributed digits of pi (to the umpteenth billion decimals), it theoretically contains copyright information in their digits.
The distribution of the copyrighted material is the infringement, but not if the data is _meant_ to produce other effects, and it is reasonable that the data is used for some other purpose _other than_ to replicate the copyrighted works.
> the training data provides value.
and so does a textbook. A student reading the book (regardless of how that book was obtained - paid or not) does not pay royalties from the knowledge obtained.
Re: Japan’s government will not enforce copyrights on data used in AI training
#210Earlier quoted context omitted.
Conflating training a model with human learning is wrong. When training a model you are deriving a function that takes some input and produces an output. The issue with copyright and licensing here is that a copy is made and reproduced numerous times when training. The model is not walking around a museum where it is an authorized viewing. It is not a being learning a skill. It is a function. The further issue is tha…
I don't fundamentally disagree with you, but what you are saying doesn't hold water. > a copy is made and reproduced numerous times when training. Casually browsing the web creates millions of copies of what are likely the same images and text that models are trained on. Computers cannot move information, they can only copy it and delete the original. Splitting hairs over the semantics of what it means to "copy" isn'…
What makes the use improper? Licenses. Terms of service. Mostly licenses though. For example, all the images on Flickr that were uploaded under Creative Commons licenses (e.g. non-commercial) have now been used in a commercial capacity by a company to create and sell a product.
Similarly, code is on Github with specific licenses with specific terms. Copilot is a derivative work of that code, the license terms of that code (e.g. GPL, non-commercial) should extend to the new function that was derived from it.
The reason I mention competition with the original is the fair use test (USA). When courts decide whether something is fair use they consider a few aspects. Two important ones are whether it is commercial, and whether it is a substitute for the original. When art models output something in the style of a living artist, it is essentially a direct substitute for that person.
Sure, I can make a shirt with Spider Man on it and give it to my brother, but if a company were to use what I made or I tried to sell it, I would expect a cease and desist from Disney.
Training the model may very well be a copyright issue. The images have been copied, they are being used. Whether that falls under fair use will likely be determined on a case by case basis in court. I do not believe closed commercial models like Copilot or Dall-e will pass a fair use test.
There is a lot of money involved here though, so we will need to wait for years before we have answers.
1. https://www.theguardian.com/technology/2012/sep/11/minnesota...