Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

51–60 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#51
post #22

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.

I've been wondering why this argument's not been sitting with me, and I think it's for the same reason that the courts have ruled that the FBI needed a warrant to put a tracker on someone's car, as opposed to following someone - the scale of action enabled is the differentiator.

A student learning from other artists is still limited in their output to human-scale - they must physically create the new thing. An AI model is not - the difference between a student learning from an artist and an AI model doing so is the AI model can flood the market with knockoffs at a magnitude the student cannot meet. Similarly, the AI model can simultaneously learn from and mimic the entirety of the art community, where the student has to focus and take time.

If this weren't capitalism - if artists weren't literally reliant on their art to eat, and if the market captured by the AI model didn't inevitably consolidate wealth - then we might be able to ignore that, but we do, and we can't ignore the economic effects when we consider scale like this.

Re: Japan’s government will not enforce copyrights on data used in AI training

#52

Earlier quoted context omitted.

No. I answered your question, do you have an answer for mine?

you can draw it but you can't [legally] sell it without permission from The Pokemon Company. And you definitely can't start a new media franchise based on your OC which combines jigglypuff with Sonic the Hedgehog.

Right, selling the drawing is what's illegal. Knowing how to draw it isn't. These models know how to draw things. Using your Pokemon ROM logic, ChatGPT should be banned because it knows how to make Pokemon-themed games.

Re: Japan’s government will not enforce copyrights on data used in AI training

#54
post #5

Is the wording accurate here? This is essentially the only source besides the untranslated article and the machine translated version sounds confusing (whether it applies to what is created by AI or what can be consumed in training).

Japanese copyright law article 30-4 states[1]:

> It is permissible to exploit a work, in ... cases ... it is not a person's purpose to personally enjoy or cause another person to enjoy ... provided, however, that this does not apply if the action would unreasonably prejudice the interests of the copyright owner ...

> i)if it is done for use in testing to develop or put into practical use technology ...

> (ii)if it is done for use in data analysis (meaning the extraction, comparison, classification, or other statistical analysis of the constituent ...

> (iii)if it is exploited in the course of computer data processing or otherwise exploited in a way that does not involve what is expressed in the work being perceived by the human senses (for works of computer programming, such exploitation excludes the execution of the work on a computer), beyond as set forth in the preceding two items.

Japanese legalese is a rather inefficient pseudo-european built on Japanese language, so I wouldn't recommend making decisions based on a blog article like this; there hasn't been too much news stories regarding this too as additional anecdotal datapoint.

1: https://www.japaneselawtranslation.go.jp/ja/laws/view/4207#j...

Re: Japan’s government will not enforce copyrights on data used in AI training

#55

So if you train an audio model on say, Eminem's voice, then write some songs and have it perform them...Would this output be legal to publish?

What's the difference between that and someone else who just happens to sound like Eminem in terms out output? As long as you don't market yourself as Eminem that should be completely legal.

Moreover, a zero shot Eminem might have a vector encoding smaller than the data required to store a fingerprint or small image.

Can you own a few numbers that represent "you"? What if someone reaches those same numbers via vocal impression or simply manually dialing and tuning some knobs and levers?

This brings up the question as to what actually makes us unique.

Re: Japan’s government will not enforce copyrights on data used in AI training

#56

So if you train an audio model on say, Eminem's voice, then write some songs and have it perform them...Would this output be legal to publish?

So I believe David Guetta of all people did this at a recent performance.

- https://finance.yahoo.com/news/dj-david-guetta-used-eminem-2...

- https://m.youtube.com/watch?v=pNP1MqmzaZg&pp=ygUWZGF2aWQgZ3V...

Re: Japan’s government will not enforce copyrights on data used in AI training

#57
post #22

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

There's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.

I'm generally against "AI same as human learning" argument but I don't think you could quite monetize recreated copyrighted arts as an art student. Van Gogh is only okay because the original artist isn't quite around.

Re: Japan’s government will not enforce copyrights on data used in AI training

#59
post #7

[flagged]

People need to understand policy making is not some binary battle between 'billionaires and commoners'. Let me give you the context behind this decision.

1. Japan is horrifically behind in software engineering compared to its neighbours, especially China. This is because of a culture that undervalues software engineers (Who don't tend to thrive in a lifetime employment/de-facto unionised environment)

2. This has started to lead into an existential threat to its entertainment industry. Genshin Impact is the most successful anime-styled game in existence (including in Japan), yet it is made by China. Why? Because China could get top tier students to work on the games, making a mobile game on a technical level that is impossible to match by Japan.

3. Now AI is out, and China is behind US, but still far ahead of Japan. Japan has two choices:

A: Let Chinese companies train their models on all of Japan's cultural outputs (All of anime/manga). Then those companies will both use the models to produce output, and sell the models back to Japanese companies. Making Japan lose by every definition. Most of the top SD anime finetunes are made by Chinese enthusiasts. So this is not a prediction, its already a reality.

B: Go all in on pro-AI policies. Giving Japanese companies legal certainty to start training their own large models. Japanese artists will still suffer losses, but at least the profits stay in Japan.

I should note Japan's artists are also not losing out that much. Anime, despite being extremely popular, is currently in complete production collapse, with even flagship shows having regular unscheduled 'breaks' because of production chaos. This collapse is because of lack of low level animators, who are so poorly paid that no one is joining the profession, and is generally outsourced anyways. AI isn't killing well paid jobs, it is replacing an already unviable job. Japan will still dominate in the higher tiers of production, and won't even need to outsource anything anymore.

Re: Japan’s government will not enforce copyrights on data used in AI training

#60
post #11

This is good news. This is basically stating that AI is inventive (which it is, in my humble opinion).

What if you overfit your model to the point of exact reproduction? Or anything in between that and what you consider inventive. Where is the line drawn.

> What if you overfit your model to the point of exact reproduction? Or anything in between that and what you consider inventive. Where is the line drawn.

The line is the same as it always has been. If you as a human publish this work then you have to respect copyright. If it is an exact replica then it likely violates copyright regardless of whether or not it was produced using AI, by copying pixels or by human hand. The test is the same regardless of how it was created.

Post reply on HN