Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

41–50 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#41

[flagged]

I mostly agree with your point but this isn’t the best analogy because while you’re allowed to learn to draw Jigglypuff, you probably could get in legal trouble for distributing or selling those drawing if it’s not fair use. I think a better analogy is using Jigglypuff drawings to learn how to draw things like that in general and then creating your own character that’s not exactly the same but uses some concepts you learnt

Re: Japan’s government will not enforce copyrights on data used in AI training

#42
post #7

[flagged]

> This is a horrible decision from Japan. Eh, I doubt it. AI is potentially something that can massively increase productivity. With a declining population in the foreseeable future, an increased productivity may well be a boost that they need.

Looks like AI is more tightly coupled to the employee (user) than the employer. AI needs prompting and supervision, that is a human skill that is tied to the employee. When the employee moves to another company, they take their AI skills with them and original company can maybe retrain the model with his data, a stop-gap measure. I think there currently is no AI that works without human involvement in critical scenarios.

Re: Japan’s government will not enforce copyrights on data used in AI training

#43
post #22

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.

AI models will make 1:1 copies of training data where artists try and avoid doing so. It’s common to obscure this copying by intentionally inserting lossy steps, but making an MP3 isn’t a new work.

It’s most obvious when large blocks of text are recreated, but the core mechanism doesn’t go away simply because you obscure the underlying output. “Extracting Training Data from Large Language Models” https://arxiv.org/abs/2012.07805

Re: Japan’s government will not enforce copyrights on data used in AI training

#44
post #22

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.

Indeed, if we don't care at all about "x is y" statements being true, they can be "applied" to reading.

To determine if an art student and DALL-E really are the same, despite their very obvious difference (one has arms and is part of a net of social relations while the other is intellectual property), will take some actual arguments which I presume you of course had planned to provide in a second comment from the start.

Re: Japan’s government will not enforce copyrights on data used in AI training

#45
post #22

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.

A student is a human and AI is not. We don’t have to apply the law equally to both regardless of how similar the method is.

Re: Japan’s government will not enforce copyrights on data used in AI training

#47

Earlier quoted context omitted.

What's the difference between that and someone else who just happens to sound like Eminem in terms out output? As long as you don't market yourself as Eminem that should be completely legal.

I suppose so. I'm just imagining these "ghost AI artists" who publish catalogues of music using the audible likeness of more prolific artists. I know that you could have always just hired an Eminem impersonator and have them lay down tracks...but this technology lets you achieve speed and scale. At least the Eminem impersonator was a real person. This is just a model learned off an artists voice.

> I'm just imagining these "ghost AI artists" who publish catalogues of music using the audible likeness of more prolific artists.

Possibly so, but who is the market for such a catalogue? I don’t see how the artist is going to lose out.

I’m reminded, though, if the episode of Mad Men where they want to get the Beatles as the soundtrack to an ad, but on finding out that the Beatles won’t do it they try to get some music that sounds similar to the Beatles. Maybe that’s the market.

Re: Japan’s government will not enforce copyrights on data used in AI training

#48
post #40

Earlier quoted context omitted.

> This is a horrible decision from Japan. Eh, I doubt it. AI is potentially something that can massively increase productivity. With a declining population in the foreseeable future, an increased productivity may well be a boost that they need.

Boost for who though? if it works out there'll be more for less. Wages won't increase, the number of jobs won't increase the price of assets will inflate. None of these are good things for 99% of us.

> Wages won't increase, the number of jobs won't increase

When new capability appears, many industries pop up. It's a new market, a new gold rush. It happened many times, with cars, electricity, air transport, computers, internet. AI will spring many applications and will create jobs in those fields.

We have been under a 260 year run of industrial revolution and 70 years of computer programming. And yet unemployment is low and IT jobs are well paid. Why do we have so many jobs? The computers are 1 million times faster now than 25 years ago, more deployed and better networked. Where is that productivity gain hiding?

If we try to be realistic, current crop AI amounts to about 1.2x productivity gain. It's a nice to have thing, but not essential yet. It's really nice. But makes errors often enough that it almost negates its advantages. Error recovery is very costly.

I foresee economic growth driven by AI, and people with AI skills being very efficient and well paid. AI shines most when it is used and then evaluated by a skilled human.

Re: Japan’s government will not enforce copyrights on data used in AI training

#50

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

No. Making a single copy for your own use is still a copyright violation. There are exceptions (fair use, nomitive use etc) but just because people are rarely sued for personal copying doesnt equate to that copying being permitted. And trademark issues, such as the other commenter generating the superman logo, are subject to a host of other rules.

> No. Making a single copy for your own use is still a copyright violation.

In some jurisdictions, perhaps, but not in all of them. There isn't one set of universal copyright law in the world. Eg in New Zealand you are allowed to make a single copy of any sound recording for your own personal use, per device that you will play the sound recording on. I'm sure there are other examples in other countries.

https://www.consumer.org.nz/articles/copyright-law

Post reply on HN