> “I swear I’ve read these instructions a hundred times but I just can’t seem to remember them”, Star complained. > Arti replied. “Let me guess: You’re rocking 102 neurals. Those won’t retain any material from Kilimanjaro. Not licensed.” > “Goddamn cheap-ass implants” grumbled Star, and handed over the instruction tablet.
Japan’s government will not enforce copyrights on data used in AI training
71–80 of 426 posts
Re: Japan’s government will not enforce copyrights on data used in AI training
#72Earlier quoted context omitted.
There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.
A student is a human and AI is not. We don’t have to apply the law equally to both regardless of how similar the method is.
Re: Japan’s government will not enforce copyrights on data used in AI training
#73Earlier quoted context omitted.
>The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. I mean, I've literally gotten the Superman logo in Stable Diffusion without even trying, so it isn't that lossy.
Superhero costume using logo. Logo inspired by the strongest gem's typical cut with first a monogram the first letter of the name for maximum size and visual clarity. Literally if you asked someone for a recognizable outline of the strongest gemstone's iconic cut you'd get the outline and the rest is an obvious path. Humans might unconsciously, or even by choice, avoid something too similar to something they already…
Re: Japan’s government will not enforce copyrights on data used in AI training
#74Earlier quoted context omitted.
There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.
Indeed, if we don't care at all about "x is y" statements being true, they can be "applied" to reading. To determine if an art student and DALL-E really are the same, despite their very obvious difference (one has arms and is part of a net of social relations while the other is intellectual property), will take some actual arguments which I presume you of course had planned to provide in a second comment from the sta…
Just a rule of thumb.
Re: Japan’s government will not enforce copyrights on data used in AI training
#75Re: Japan’s government will not enforce copyrights on data used in AI training
#76I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…
There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.
The AI is not a human, but what you are doing is the same thing, if you claim the output as your work because you wrote the prompt.
Re: Japan’s government will not enforce copyrights on data used in AI training
#77Earlier quoted context omitted.
There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.
AI models will make 1:1 copies of training data where artists try and avoid doing so. It’s common to obscure this copying by intentionally inserting lossy steps, but making an MP3 isn’t a new work. It’s most obvious when large blocks of text are recreated, but the core mechanism doesn’t go away simply because you obscure the underlying output. “Extracting Training Data from Large Language Models” https://arxiv.org/ab…
In general I don't think this is the case, assuming you mean generations output from popular text-to-image models. (edit: replied before their comment was edited to include the part on text generation models)
For DALL-E 2: I've never seen anyone able to provide a link of supposed copying. Even if you specifically ask it for some prominent work, you get a rendition not particularly closer than what a human artist could do: https://i.imgur.com/TEXXZ4a.png
For Stable Diffusion: it's true that Google did manage, by generating hundreds of millions of images using captions of the most-duped training images and attempting techniques like selecting by CLIP embeddings, to get 109 "near-copies of training examples". But I'd speculate, particularly if you're using the model normally and not peeking inside to intentionally try to get it to regurgitate, that this is still probably lower than the human baseline rate of intentional/accidental copying. It does at least seem lower than the intra-training-set rate: https://i.imgur.com/zOiTIxF.png (though many may be properly-authorized derivative works)
Re: Japan’s government will not enforce copyrights on data used in AI training
#78So it's kind of disgusting to see this relaxation, when it suits some corporate or national interests.
https://en.wikipedia.org/wiki/File_sharing_in_Japan
"Unlike most other countries, filesharing copyrighted content is not just a civil offense, but a criminal one, with penalties of up to ten years for uploading and penalties of up to two years for downloading."
"There is also a high level of Internet service provider cooperation."
Re: Japan’s government will not enforce copyrights on data used in AI training
#79Earlier quoted context omitted.
I'm generally against "AI same as human learning" argument but I don't think you could quite monetize recreated copyrighted arts as an art student. Van Gogh is only okay because the original artist isn't quite around.
Can anyone monetize Van Gogh regardless? If a human or AI reproduces a Van Gogh painting or derivative, it's not worth anything on the market. Only original pieces, by a human artist, has real value. A Van Gogh painting is worth millions of dollars only because it was created by Van Gogh. A reproduction is approximately worth the paper it's printed on.
Re: Japan’s government will not enforce copyrights on data used in AI training
#80I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…
There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.
When training a model you are deriving a function that takes some input and produces an output. The issue with copyright and licensing here is that a copy is made and reproduced numerous times when training.
The model is not walking around a museum where it is an authorized viewing. It is not a being learning a skill. It is a function.
The further issue is that it may output material that competes with the original. So you may have copyright violation in distribution of the dataset or a model's output.