Earlier quoted context omitted.
Not just crypto bros, this kind of thing is rife in politics too. Brexit is full of it. People pick one article about one minor thing in one niche area of the economy and use it to 'prove' their entire agenda.
Not just politics. I remember this being a realisation as a teenager, noticing that if you bring five reasons, the person you're talking to will refute a random one in a funny way and now the audience will decide you were wrong. Danny was the person who was absolutely the best at this. I should have written one of them down, as I can't even reproduce it but he'd use some logical fallacy to make his case which is, for…
Japan’s government will not enforce copyrights on data used in AI training
351–360 of 426 posts
Re: Japan’s government will not enforce copyrights on data used in AI training
#352Earlier quoted context omitted.
If you search "Superman Logo" you find actual copies of the Superman logo which are served from Google's cache. If you ask a VFX artist to create the "Superman Logo" with Photoshop they'll do an excellent job. The first one isn't copyright violation because it is fair use. The second maybe if it is redistributed but we don't ban the use of photoshop by artists because they can choose to reproduce copyright things wit…
I agree, and I honestly think that a big part of the issue with AI image generation is people just really have a hard time conceiving of a technology that can make such accurate images from a relatively small model like this. "It must have a copy" - but van Gogh didn't make paintings of hot rods or whatever, and you can't copyright style or technique.
Re: Japan’s government will not enforce copyrights on data used in AI training
#353Re: Japan’s government will not enforce copyrights on data used in AI training
#354Earlier quoted context omitted.
> If you are going to use someone else's work in order to make something that you are going to profit off of, I believe that original author should be compensated. And should also be able to decide they don't want their work used in that way. > Note that I'm not talking about what existing copyright law says; I'm talking about how I believe we should be regulating this new facet of the industry. Is it really new? Hum…
> You can make private works if you want to keep control of them, but at some point the public deserves to share and rework the things that have been pushed into the public consciousness. There's already a licensing framework for artists doing this - should they wish to. It's called Creative Commons, and allows a pretty fine distinction of rights from public domain to free for personal use not commercial, and everyth…
Re: Japan’s government will not enforce copyrights on data used in AI training
#355If I compress an artist's painting into a jpeg and rehost part of it for individual t-shirt designs I am committing a crime. If I compress an artist's painting into a model & rehost what's essentially a highly flexible complete version of their painting for infinite, perpetual use of any kind ... I'm not committing a crime?
>compress an artist's painting into a model That's not how image models work.
That the compression technology relies on parameterizing the copyrighted material, and as a result can produce hallucinations remixing the copyrighted material, is super cool but doesn't change that this is compression at rest.
All the hullabaloo about artificial intelligence is science fiction laundering a (cool new) compression algorithm.
Copyright law shouldn't apply any differently to a LLM as it does to gzip.
Re: Japan’s government will not enforce copyrights on data used in AI training
#356Earlier quoted context omitted.
> Humans have always learnt by studying what's out there already. Our whole culture is built on what's been done and published before Are you implying that educators should not be compensated or credited? Because that is not how it works in the real world.
If I read a lot of fantasy books as a kid, then start writing my own fantasy book, should I have to pay royalties to the authors of the books I read?
Most jobs require a degree or certification of some sort.
Re: Japan’s government will not enforce copyrights on data used in AI training
#357This article is an example of emerging AI-bro tactics that completely mirrors crypto-bro tactics: they pick any piece of news and reinterpret it to fit an agenda. While the article is in English, the link to source is in Japanese. The only external source I found suggests the discussion is about promoting open data and open science from research institutions [1] [1] https://asianews.network/japan-to-promote-use-of-ge…
The Japanese article does explicitly state it if you run it through a translator, and also this is from May 11 > Additionally, the group raised other issues that Article 30-4 of the Copyright Law, which permits the use of a copyrighted work for machine learning, does not include procedures for gaining permission in advance from copyright holders. The article permits the use of copyrighted material such as text and im…
> まずAIによる情報解析についての我が国の法制度(著作権法)について確認したところ、我が国において、非営利目的であろうと、営利目的であろうと、複製以外の行為であろうと、違法サイトなどから取得したコンテンツであろうと、方法を問わず情報解析のための作品利用はできると永岡大臣が明言しました。
> Confirming the legal system (copyright law) wrt. data analysis by AI in our country, Minister Nagaoka clearly stated that in our country, whether for non-profit purposes or for profit purposes, whether an act other than reproduction, or whether the content is obtained from illegal sites, one can use works for information analysis regardless of the method.
(translated with ChatGPT-4 and then cleaned up)
The source is the one from the article: https://go2senkyo.com/seijika/122181/posts/685617
Re: Japan’s government will not enforce copyrights on data used in AI training
#358Earlier quoted context omitted.
> The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. Lossy or not, the training data provides value . If all the various someones had not spent time making all the stuff that ends up as training data, then the model it trains would not exist. If you are going to use someone else's work in order to make something that you are going to…
> Lossy or not, the training data provides value. If we ignore the issue of machine learning for now; It's not the job of copyright to prevent people extracting value from a copyrighted work. If it was, then it would be possible for copyright holders to launch lawsuits that block entities from using the knowledge that was published in copyrighted reference material. Or the rights holder of a cookbook would be able to…
I would personally also argue that machine learning model is roughly equivalent to a compression algorithm. Converting a 4k video to a 420p video is just as lossy as feeding a learning model from a 4k video and ask it to reproduce the 420p video. It has nothing in common with how a human brain is consuming content or learns information. No person can produce a 420p video just by consuming a 4k video, nor can any machine learning model gain the emotional constructs and social contexts that human brains get from learning.
Re: Japan’s government will not enforce copyrights on data used in AI training
#359Earlier quoted context omitted.
Could be argued, sure. If you have to already have access to the copyrighted images to find them in the model, the argument seems weak. A sufficiently advanced model could, in theory, generate any image. You could then, again in theory, find an embedding for any image. Does said model then infringe on all copyrighted images? A program that creates Fourier epicycle drawings could be given input that causes trademarked…
> If you have to already have access to the copyrighted images to find them in the model, the argument seems weak. That makes no sense. The copyright holder has access to their own inventions, of course. That's the standard in any copyright claim. > A sufficiently advanced model could, in theory, generate any image. You could then, again in theory, find an embedding for any image. Does said model then infringe on all…
No, not talking about the copyright holder, I'm talking about the hypothetical individual(s) creating infringing copies. If those people need to already have a copy of the image to extract a copy of the image from the generative image model, then I'm saying the argument that the model itself is infringing seems weak. Or it's at least not an open-and-shut case.
>> A sufficiently advanced model could, in theory, generate any image. You could then, again in theory, find an embedding for any image. Does said model then infringe on all copyrighted images?
> Without the slightest doubt.
I'll continue to argue otherwise. This proposed model is not a compressed archive that reproduces a set of infringing works when decompressed. Instead, you already have to have a copy of an image to find an embedding. Otherwise, the chances of the model spitting out copies of infringing works is exceedingly improbable. (A program that outputs random noise also has a vastly improbable chance of spitting out a copyrighted work, but that's hardly keeping copyright holders up a night.) Furthermore, in being able to produce any image, the model is not going to contain every image, and provided a copyrighted image produced after the creation of the model, you could still find an embedding. From a copyright perspective, suing the creator of this model would be like suing someone over distributing an image of random noise, claiming that because you can find an "embedding" which produces your copyrighted work (really just the difference between the two images), the noise is infringing.
Now, if you want to sue someone for distributing an embedding into this model for infringement, that's another matter entirely. That makes perfect sense.
In reality, I acknowledge that models like Stable Diffusion are going to be a bit more muddy. There definitely is some overfitting going on, so some images are literally present. However, it's a case-by-case thing. Given the requirements (framed as an "attack" no less) for extracting those images, a particular release of SD might or might not be found to infringe. Other models, with better training and better datasets, could avoid the overfitting problem.
> The method of storing the information is pretty much irrelevant to copyright. Your link argument has been tried by pirates and it's not working too well, although it depends on the country and legislation.
Unless you agree that Pi itself is a copyright violation, I think you misunderstand. I'm not making the same "link" argument made by pirates. The index and length needed to find copyrighted embeddings in Pi is just a different encoding for the same data, similar to a compressed version of the same data, though I'm sure the Pi embedding would in fact tend to be absurdly larger than the original. Again, I'm saying that Pi isn't infringing here, but the Pi-rates with their "links" would be where the infringement happens.
Re: Japan’s government will not enforce copyrights on data used in AI training
#360Earlier quoted context omitted.
>compress an artist's painting into a model That's not how image models work.
LLM training is a compression algorithm, LLM weights are a compressed dataset of the training materials, and executing LLMs is accessing the compressed source material. That the compression technology relies on parameterizing the copyrighted material, and as a result can produce hallucinations remixing the copyrighted material, is super cool but doesn't change that this is compression at rest. All the hullabaloo abou…