Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

61–70 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#61
post #28

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

>The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. I mean, I've literally gotten the Superman logo in Stable Diffusion without even trying, so it isn't that lossy.

Superhero costume using logo. Logo inspired by the strongest gem's typical cut with first a monogram the first letter of the name for maximum size and visual clarity.

Literally if you asked someone for a recognizable outline of the strongest gemstone's iconic cut you'd get the outline and the rest is an obvious path. Humans might unconsciously, or even by choice, avoid something too similar to something they already know.

Superman's costume also uses vibrant colors. The red / blue pairing is used extensively across many logos and visual representations for the high contrast of two vibrant colors.

As I try to imagine an older child or young adult somehow raised in an environment like pop culture but through some twist absolutely unexposed to Superman or any related concepts, it isn't that far of a stretch to imagine independent invention of a strikingly similar idea. Maybe not as a first draft but in exploring a range of possible powers and automatic logos. E.G. as in the range of an LLM backed character creator for a superhero game, and then aneling the results though simulated effectiveness / fitness of hero powers, logo design, etc.

Everyone wants to think they're a special snowflake and that what they create is somehow unique as well. However we're all drawing on a huge pool of common culture to synthesize expressions which fulfill a set of constraints prescribed by the culture and the culture's influence on the individual and the moment being experienced.

In the case of Superman that's even arguably a description of the archetype. They are literally a super man. Clark Kent however, that's a little more unique and probably a Trade Mark (consumer commercial use protection) as long as such a registration is maintained.

Re: Japan’s government will not enforce copyrights on data used in AI training

#62
post #11

This is good news. This is basically stating that AI is inventive (which it is, in my humble opinion).

What if you overfit your model to the point of exact reproduction? Or anything in between that and what you consider inventive. Where is the line drawn.

I don’t think the tool by which a derivative work is created matters. The author should not distribute, remix, re-work, adapt, release, perform or synchronise it if they do not have the rights. This is already the case for all other technologies and tools.

For example, if you rewrite a popular book in a word processor from scratch, it is not the responsibility of your word processor to not let you do it. Or the government to regulate word processors so that they are incapable of doing it. It is your responsibility to not distribute what you made.

If you record your screen watching a movie and distribute it, the accountability for this falls on you, not the tool made for screen recording.

And if you use generative AI to produce trademarked or copyrighted works, you should be accountable, not AI.

Generative AI will always be able to produce copyrighted works if steered enough. Even if it wasn’t trained directly on them. A very simple experiment proves it - you can paste copyrighted material into a ChatGPT prompt and ask it to repeat it. Or you can describe a copyrighted work really well for MidJourney and it will produce results with high likeness. This does not require that the models be trained on copyrighted works.

Besides, copyright infringement isn’t nearly the worst crime AI can be used in. Serious impersonation, forgeries and fraud are also made easier with AI. Why is copyright different?

It should all be treated the same - the user should be held accountable for their actions. Not the tool or the tool’s makers.

Re: Japan’s government will not enforce copyrights on data used in AI training

#63

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

No. Making a single copy for your own use is still a copyright violation. There are exceptions (fair use, nomitive use etc) but just because people are rarely sued for personal copying doesnt equate to that copying being permitted. And trademark issues, such as the other commenter generating the superman logo, are subject to a host of other rules.

Training a model isn’t making a copy for your own use, it’s not making a copy at all. It’s converting the original media into a statistical aggregate combined with a lot of other stuff. There’s no copy of the original, even if it’s able to produce a similar product to the original. That’s the specific thing - the aggregation and the lack of direct reproduction in any form is fundamentally not reproducing or copying the material. The fact it can be induced to produce copyright material, as you can induce a Xerox to reproduce copyright material, doesn’t make the original model or its training a violation of copyright. If it’s sole purpose was reproduction and distribution of the material or if it carried a copy of the original around and produced it on demand, that would be a different story. But it’s not doing any of that, not even remotely. All this said, it’s a highly dynamic area - it depends on the local law, the media and medium, and the question hasn’t been fully explored. I’m wagering though when it comes down to it, the model isn’t violating copyright for these reasons, but you can certainly violate copyrights using a model.

Re: Japan’s government will not enforce copyrights on data used in AI training

#64
post #28

Earlier quoted context omitted.

>The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. I mean, I've literally gotten the Superman logo in Stable Diffusion without even trying, so it isn't that lossy.

And if you use it, you’re violating copyright. But you will find no copy of the logo in the model data. The model is way too small to contain its training imagery from an information theoretic point of view.

> But you will find no copy of the logo in the model data.

You wont find a copy of a plaintext in a cyphertext. But you can still extract the plaintext from the cyphertext.

Re: Japan’s government will not enforce copyrights on data used in AI training

#65
> “I swear I’ve read these instructions a hundred times but I just can’t seem to remember them”, Star complained.

> Arti replied. “Let me guess: You’re rocking 102 neurals. Those won’t retain any material from Kilimanjaro. Not licensed.”

> “Goddamn cheap-ass implants” grumbled Star, and handed over the instruction tablet.

Re: Japan’s government will not enforce copyrights on data used in AI training

#66
post #22

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.

One day we'll have a thread about AI where someone doesn't use the "machines deserve the same rights as people" non-argument. But this isn't that thread.

Re: Japan’s government will not enforce copyrights on data used in AI training

#67
post #22

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.

> There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training.

A few notable differences:

1. Scale: a single art student can't view millions of works in a week.

2. Duplication: a single art student's brain can't be cloned or downloaded into another art student's brain.

3. Speed: a single art student cannot draw or paint thousands of images in a week.

4. Ownership: the software is likely owned and controlled by a large corporation, while the art student (hopefully) isn't.

Re: Japan’s government will not enforce copyrights on data used in AI training

#68
post #7

[flagged]

Initially I was skeptical too but maybe Japan really can increase its GDP by 50% via AI-generated anime? Would love to see the government report that backs this claim up though ...

Joke on them, you are already getting unsolicited anime content when using Stable Diffusion.

Re: Japan’s government will not enforce copyrights on data used in AI training

#70
post #57
post #22

Earlier quoted context omitted.

There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.

I'm generally against "AI same as human learning" argument but I don't think you could quite monetize recreated copyrighted arts as an art student. Van Gogh is only okay because the original artist isn't quite around.

Can anyone monetize Van Gogh regardless?

If a human or AI reproduces a Van Gogh painting or derivative, it's not worth anything on the market.

Only original pieces, by a human artist, has real value. A Van Gogh painting is worth millions of dollars only because it was created by Van Gogh. A reproduction is approximately worth the paper it's printed on.

Post reply on HN