Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

281–290 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#281
post #276

I have changed my mind on this recently. Sure, exploiting free software to the benefit of private corps is bad but if the law would allow us to train an open source net on LibGen (with all the copyrighted books and papers) and then to distribute the weights legally, I am all for that.

Copyright has always been a pretty dumb concept brought upon by the issue that "thinkers" wanted a bigger piece of the pie.

Don't get me wrong, I can totally understand their reason: how can an author make a living if a printing shop could just start producing copies of their book (that's the context the law was passed in)... But it's arguably a way too blunt instrument which gives the copyright holder a disproportionate amount of power vs someone producing physical goods.

I don't claim to have an answer to this problem and it's likely another instance of having a flawed system that works well enough that the upsides outweigh the downsides... Like so many other things in our society such as capitalism and representative democracy

Re: Japan’s government will not enforce copyrights on data used in AI training

#282
post #138

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

> The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. Lossy or not, the training data provides value . If all the various someones had not spent time making all the stuff that ends up as training data, then the model it trains would not exist. If you are going to use someone else's work in order to make something that you are going to…

Value derived from a work was never a component of protected IP.

If you write a song which is so inspirational that it influences the way a listener thinks or even inspires them to make similar sounding songs, you don't have any claim to that value.

Re: Japan’s government will not enforce copyrights on data used in AI training

#284

Earlier quoted context omitted.

I hope smarter minds than myself will somehow figure out a way to square this circle and get heavy restrictions on commercial usage without killing the technology outright. As far as AI rap goes: Notorious BIG is (by far) my favorite rapper of all time, yet he died in his early 20s with barely any unreleased songs in the vault. For me, as an aging superfan, AI deepfakes have me feeling 15 again - like Biggie is still…

There's a saying in physics that the field "advances one funeral at a time". I don't understand the impulse to bring back artists of the past like this. At best, it's a hollow pantomime, and at worst it can be tremendously offensive too their memory. Moreover, how will the medium move forward if every new artist has to now compete for mindshare with every artist that lived and died since 1950 or so? Not just their le…

I personally agree that AI art is likely a cultural cul-de-sac, but I don’t think our current society has a framework that can exclude it from the art world without inflicting an Emperor Has No Clothes situation on several decades of post-modernist art. So there will be much gnashing of teeth, but ultimately the question posited by AI (“what is Art?”) is too destructive/inconvenient to be answered and the hubbub will eventually go away as AI generation becomes normalized.

Doomerism aside, I think the focus should be on stopping wide scale commercialization of AI generation being packaged as art. I greatly enjoy the dead artist pantomimes, but I certainly don’t want Diddy magicking up a new Biggie single through AI necromancy.

Re: Japan’s government will not enforce copyrights on data used in AI training

#285
post #152

Japan also ranks 3rd (behind the USA & India, with larger populations) in ChatGPT usage: https://www.demandsage.com/chatgpt-statistics/ There's also been discussion of their government using ChatGPT to reduce red tape: https://www.bloomberg.com/news/articles/2023-04-18/japan-gov... It's cool to see Japan and Japanese culture taking techno-optimist stances on AI.

Which I find bizarre given how backwards Japan is in the adoption of other technologies. Eg their continued reliance on paper records and fax machines.

I've heard this FAX meme since like early 2000, but I have yet to encounter one in my 4 years living here, and most Japanese people I joke this too just gets as perplexed by the joke.

I wonder, where do people find those FAX machine? Have I lived in a tech/startup bubble in Tokyo and missed it? I didn't even see it at the local Ward office in the suburb.

Some stuff are still old-school (hanko etc) but it to me seem like the fax meme have outlived the reality.

Last time I heard fax machine was a German exchange student in Tokyo, AFTER she returned to Germany and had to get some paperwork at the ward office, in Germany

Re: Japan’s government will not enforce copyrights on data used in AI training

#286
This article is an example of emerging AI-bro tactics that completely mirrors crypto-bro tactics: they pick any piece of news and reinterpret it to fit an agenda.

While the article is in English, the link to source is in Japanese. The only external source I found suggests the discussion is about promoting open data and open science from research institutions [1]

[1] https://asianews.network/japan-to-promote-use-of-generative-...

Re: Japan’s government will not enforce copyrights on data used in AI training

#287

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

Some licenses, like CC, have variants that prohibit production of derivatives, or prohibit commercialization, or require licenses or same requirements to be preserved in derivative works. Sweet to see people "warm up" to say stuff like 'well it's just a derived work'. Can we get tech to actually respect the licenses of used works next? It's something that's been asked for all along. Without just going, 'well, it's all fair use' - 'so we'll just ignore all licenses and won't even try to detect or respect licenses and whatever requirements they have'. Sure, it may be "unenforceable", but if tech keeps saying "fuck you" to the creators, and "fuck you specifically to the licenses, we won't even look at them or process them" - creators will keep saying 'well fuck you too' right back.

otherwise, it's just talk with no follow through. just dropping words like 'fair use' as a 'get away from liability card', without actually engaging with intellectual property concepts. just to get to use works without respect to artist's will (as it could be expressed in a license), or sometimes even a mention, let alone "compensation" or other "consequence".

Re: Japan’s government will not enforce copyrights on data used in AI training

#288

Earlier quoted context omitted.

Eh, I disagree. Copyright laws are mostly bullshit anyway, and only tend to favor capital holders, who tend to buy up all the copyright they need. I would gladly see copyright rendered useless. The peasantry hardly benefits from it anyway.

"AIs can ignore copyright" is the absolute ultimate in blank check for capital holders. Waiving copyright to solve the problem of systems favoring them is like deciding to jump because you're afraid of heights. Copyright law has done a huge amount to reward creators -- I know even local-tier artists and musicians without the support of large capital who make a good chunk of their living through sales supported by it.…

How do I benefit from 90 years long copyright terms?

Re: Japan’s government will not enforce copyrights on data used in AI training

#289
post #287

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

Some licenses, like CC, have variants that prohibit production of derivatives, or prohibit commercialization, or require licenses or same requirements to be preserved in derivative works. Sweet to see people "warm up" to say stuff like 'well it's just a derived work'. Can we get tech to actually respect the licenses of used works next? It's something that's been asked for all along. Without just going, 'well, it's al…

Did you just argue against fair use in general? Everything you wrote applies to it as well. Copyright has limits, it's not like right holders get to determine what those are. It's a balance of rights of creators and users.

Re: Japan’s government will not enforce copyrights on data used in AI training

#290
post #138

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

> The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. Lossy or not, the training data provides value . If all the various someones had not spent time making all the stuff that ends up as training data, then the model it trains would not exist. If you are going to use someone else's work in order to make something that you are going to…

> Lossy or not, the training data provides value.

If we ignore the issue of machine learning for now; It's not the job of copyright to prevent people extracting value from a copyrighted work.

If it was, then it would be possible for copyright holders to launch lawsuits that block entities from using the knowledge that was published in copyrighted reference material. Or the rights holder of a cookbook would be able to block people from making the recipe.

We have other sections of law that provide protection along these lines. Patents give their holders a monopoly to extract value from the given invention, a much stronger protection that copyright law. But in exchange Patents are limited to 21 years, much be publicly documented and only certain types of things are covered.

Trade secret laws can be used to protect recipes, formulas and other processes, but only as long as the holder makes reasonable efforts to keep it secret. The owner can't have it both ways, have the IP publicly known and protected by trade secret law.

The only purpose of copyright is to give the holder a monopoly over the reproduction of a work. The definition of reproduction might be quite wide these days: A performance of a work is a reproduction, a cover of a song is a reproduction, distribution is reproduction in the modern age of computers... etc, but the limit of copyright law is reproduction.

Coming back to machine learning, there is a somewhat open question if training counts as a form of reproduction (well, outside of Japan). But we can't use your proposed "extraction of value" metric as a way to decide that.

Personally, I would argue that training a machine learning model is roughly equivalent to a human brain consuming copyrighted works, and should be treated the same in law.

The fact that a machine learning model is (sometimes) capable of recreating a copyrighted work later shouldn't be held against them, as a human brain is also fully capable of recreating copyrighted works from memory. From a legal perspective, recreating a copyrighted work from memory will not save a human from a copyright infringement lawsuit, it's simply not a defence. The copyright infringement happens when the work is recreated.

Post reply on HN