Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

331–340 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#331
post #126
post #88

Earlier quoted context omitted.

The more degrees of freedom the less likely independent creation rather than copping occurred. LLM’s recreating training material causes real issues such as Google’s dealing with PII leaks: https://ai.googleblog.com/2020/12/privacy-considerations-in-... If one prompts the GPT-2 language model with the prefix “East Stroudsburg Stroudsburg...”, it will autocomplete a long block of text that contains the full name, phon…

Privacy, where there's a problem if some original data can be inferred/made out (even using a white box attack against the model), is a higher bar than whether an image generator avoids copyright-infringing output under non-adversarial usage. Additionally, compared to image data, text is more prone to exact matches due to lower dimensionality and usually training with less data per parameter. While it's still a topic…

The privacy issue isn’t just about the data being available as people’s names, addresses, and phone numbers are generally available. The issue is if they show up as part of some meme chat and then you as the LLM creator get sued because people start harassing them.

In terms of copyright infringement the bar is quite low, and copying is a basic part of how these algorithms work. This may or may not be an issue for you personally but it is a large land mine for commercial use especially if you’re independently creating one of these systems.

Re: Japan’s government will not enforce copyrights on data used in AI training

#332

Earlier quoted context omitted.

Eh, I disagree. Copyright laws are mostly bullshit anyway, and only tend to favor capital holders, who tend to buy up all the copyright they need. I would gladly see copyright rendered useless. The peasantry hardly benefits from it anyway.

I don't think you understand what's being proposed here. I expect that, as today, the AI output will absolutely be copyrighted and strictly protected against any attempt to copy it by others. However these protections won't apply to the inputs a company will use to train the AI. ...and oh yeah, if the Sam Altmans of the world get their way you'll need a license from the government to run your own AI model. We might g…

> I don't think you understand what's being proposed here. I expect that, as today, the AI output will absolutely be copyrighted and strictly protected against any attempt to copy it by others. However these protections won't apply to the inputs a company will use to train the AI.

And the output of that will not be copyrightable to train AI, whose output will be copyrightable, but not to train more AI, ad infinitum.

Copyright is the problem here, not AI.

> ...and oh yeah, if the Sam Altmans of the world get their way you'll need a license from the government to run your own AI model.

That may be impossible to enforce, as models leak into the world, and you can run then offline in a sufficiently powerful machine.

Let's hope that becomes the case.

Re: Japan’s government will not enforce copyrights on data used in AI training

#333
post #229
post #138

Earlier quoted context omitted.

> The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. Lossy or not, the training data provides value . If all the various someones had not spent time making all the stuff that ends up as training data, then the model it trains would not exist. If you are going to use someone else's work in order to make something that you are going to…

> If you are going to use someone else's work in order to make something that you are going to profit off of, I believe that original author should be compensated. And should also be able to decide they don't want their work used in that way. > Note that I'm not talking about what existing copyright law says; I'm talking about how I believe we should be regulating this new facet of the industry. Is it really new? Hum…

> You can make private works if you want to keep control of them, but at some point the public deserves to share and rework the things that have been pushed into the public consciousness.

There's already a licensing framework for artists doing this - should they wish to. It's called Creative Commons, and allows a pretty fine distinction of rights from public domain to free for personal use not commercial, and everything in between. https://creativecommons.org/

I agree our shared culture in some sense ultimately owns the productions of the culture. But that's wildly different to (and in some senses the opposite of) letting private companies enclose, privatise and sell those products back to us. As for example Disney has done over and over again, taking myths transcribed by the Brothers Grimm, or classic novels now in the public domain, 'reinterpreting' them and viciously enforcing copyright on these new interpretations.

The entire point of copyright law is to allow the Bach's of this world to profit from their work - without having to die in poverty and obscurity, as so many artists and musicians have historically, even while others have profited from their work at scale.

> Humans have always learnt by studying what's out there already.

Finally - as other commentators have noted, there's really no similarity at all between a human at human pace studying and integrating understanding of a piece or genre of art, and an AI training to replicate that work at scale as perfectly as possible. A much better comparator would be the Chinese factory 'villages' that reproduce paintings at scale for the commercial market, without creativity or 'art' being a part of the process. But even that is a poor analogy, since individual humans mediate the process. A really good analogy would be a giant food corporation like Nestle somehow scanning the product of a restaurant and then offering that chefs unique dishes that had taken years to invent for nearly free - using the same name and benefiting from the association.

Re: Japan’s government will not enforce copyrights on data used in AI training

#334
post #276

I have changed my mind on this recently. Sure, exploiting free software to the benefit of private corps is bad but if the law would allow us to train an open source net on LibGen (with all the copyrighted books and papers) and then to distribute the weights legally, I am all for that.

Copyright has always been a pretty dumb concept brought upon by the issue that "thinkers" wanted a bigger piece of the pie. Don't get me wrong, I can totally understand their reason: how can an author make a living if a printing shop could just start producing copies of their book (that's the context the law was passed in)... But it's arguably a way too blunt instrument which gives the copyright holder a disproportio…

Take a look at https://kottke.org/17/12/unlocking-the-commons-or-the-psycho... and mutualism, patronage, crowdfunding, bounties and commissions, which seem to be good alternative models for post-scarce goods such as digital data.

Re: Japan’s government will not enforce copyrights on data used in AI training

#335
post #50

Earlier quoted context omitted.

No. Making a single copy for your own use is still a copyright violation. There are exceptions (fair use, nomitive use etc) but just because people are rarely sued for personal copying doesnt equate to that copying being permitted. And trademark issues, such as the other commenter generating the superman logo, are subject to a host of other rules.

> No. Making a single copy for your own use is still a copyright violation. In some jurisdictions, perhaps, but not in all of them. There isn't one set of universal copyright law in the world. Eg in New Zealand you are allowed to make a single copy of any sound recording for your own personal use, per device that you will play the sound recording on. I'm sure there are other examples in other countries. https://www.c…

This is the same in the UK (and not only for sound). If you own the copy, you can make personal copies. You can't share them, and you have to own the original.

Re: Japan’s government will not enforce copyrights on data used in AI training

#336
post #288

Earlier quoted context omitted.

How do I benefit from 90 years long copyright terms?

The same way you benefit from copyright being valuable to creators at all combined with the reason why options with longer exercise horizons are more valuable. Honestly I'm not sure 90 years or whatever is the optimal time and I'm certainly willing to sign on to discussion with thoughtful people or even political movements about whether term length reform should be included in those "when it doesn't [work]" considera…

I am not convinced. Creators have multiple avenues to profit from their creations beyond copyright.

Either way, this is orthogonal to the original point. Copyright being undermined by generative AI may be a positive outcome for society as a whole, as it is a powerful creative tool. Some things that might take a large team of people to create will perhaps be more accessible to solo creators or smaller teams.

> But also I'm really tired of having that conversation with people who aren't thoughtful and ask questions like that as if it's some kind of insightful point about the copyright model in general when the truth is that it's a just a parameter.

Then don't have that conversation. No one is forcing you.

Re: Japan’s government will not enforce copyrights on data used in AI training

#337
post #89

Earlier quoted context omitted.

I strongly agree with this. There's a distinction between "learning from" and "copying". "Learning from" is a transformative process that distills from the observation. This distillation can be as simple as indexing for a search engine, or as complex as a deep neural network. Simply because a neural network can create something that is a copyright violation doesn't mean the training process itself it. A human can see…

> There's a distinction between "learning from" and "copying". Neural nets can memorize their training data. Generally that isn't what you want, and you strive to eliminate it. However, it could instead be encouraged to happen if someone wanted to exploit this law in order to abuse copyrights.

Yes, and as GP suggested, going on to distribute copies would be copyright infringement. That doesn't imply that it's an infringement to train the neural net.

Re: Japan’s government will not enforce copyrights on data used in AI training

#338
post #138

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

> The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. Lossy or not, the training data provides value . If all the various someones had not spent time making all the stuff that ends up as training data, then the model it trains would not exist. If you are going to use someone else's work in order to make something that you are going to…

> Lossy or not, the training data provides value. If all the various someones had not spent time making all the stuff that ends up as training data, then the model it trains would not exist.

I don’t think deriving value is the measure for infringement. The artist derived value from society and from countless inspirations, should they be compensated?

Copyright is the right to distribute your content, not to control it in every way and derive maximum value possible. The purpose of copyright is to incentivize creators and they already get, in the US, 70+ years of monopoly preventing anyone else from selling copies.

This seems like more than enough incentive for creators to create, as evidenced by massive copyrighted material created at a rate far above population growth (more content is created and more money is made via copyright than ever before).

Re: Japan’s government will not enforce copyrights on data used in AI training

#339

Earlier quoted context omitted.

If I read a lot of fantasy books as a kid, then start writing my own fantasy book, should I have to pay royalties to the authors of the books I read?

Does your ability to write fantasy books absolutely depend on having read those fantasy books as a kid? Was gaining the ability to write your own fantasy books and profit from them your only motivation to read those fantasy books? After gaining the ability to write fantasy books thanks to having read them, can you now produce fantasy books at a qualitatively different speed, scale, and conditions than any of the auth…

They already paid when they bought the books. Why would they need to pay more?

Re: Japan’s government will not enforce copyrights on data used in AI training

#340

Earlier quoted context omitted.

How can you be certain that the content being generated is non-infringing?

The same way you do with any other content you generate in other ways.

Well, when I pick up a pencil and make a drawing, I have a lot of agency over what is created.

The whole point of these generative models is that I have less agency over exactly what gets created - it takes my prompt and does the rest.

Post reply on HN