Live data from Hacker News

Japan Goes All In: Copyright Doesn't Apply to AI Training

biia.com

81–90 of 183 posts

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#81
It is not that hard to train on public-domain content only or content that has the permission of the creators. Stability knew that they would get into trouble if they dared to evade copyright laws in the music industry with StableAudio. [0]

We'll see how Sony Music Japan, Universal Music Group Japan and the rest think about this latest iteration of regulatory arbitrage with AI.

[0] https://stability.ai/research/stable-audio-efficient-timing-...

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#82

I think it’s a weak argument anyway. A pencil can be used to recreate copyrighted works. Why should a LLM be different? The legal responsibility always has been and should continue to be on the person publishing or making available the work. And at any rate, a country banning this tech will be missing the revolution. Protecting the buggy whip manufacturers and all that.

What's the difference between picking a cherry tomato from your streetside garden and eating it, and driving a combine harvester over your lawn and taking everything?

Legally, the scale of the punishment / fine, but not the legality of the base act. Also, the latter involves trespassing and probably property damage.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#83
Policymakers are making a tradeoff between privacy and economic value. When governments see the potential for creating large tech companies and loose copyright policy can give them an edge, they must decide to choose between privacy or economic value.

South Korea has mirrored Japan's policy: https://metanews.com/south-korean-government-says-no-copyrig...

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#84

I think it’s a weak argument anyway. A pencil can be used to recreate copyrighted works. Why should a LLM be different? The legal responsibility always has been and should continue to be on the person publishing or making available the work. And at any rate, a country banning this tech will be missing the revolution. Protecting the buggy whip manufacturers and all that.

What’s the work? Would you say it ought to be legal to distribute a really good prompt to generate a copyright character? What about an embedding?

I would say it ought to be as legal as distributing a "how to draw mickey mouse" tutorial or a "how to sing taylor swift song" video or perhaps even a "how to make a twitter clone" tutorial.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#85
post #41

There is a distinction to be made between inputs and outputs when it comes to AI and copyright. Most people focus on the former, and discuss whether you can train an LLM on copyrighted works or not, but ultimately the issues really only manifest in the latter. An AI training itself on a million newspaper articles can be declared legal, sure, but what happens when it also starts spitting out the same articles with nea…

> This is what the crux of NYT's lawsuit is about, and making laws about AI training isn't going to make a difference to that. The claims of NYT are more than just about training. They're claiming that as part of the ChatGPT software; it _looks up_ stuff in a database of articles. That is beyond _training_. If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think…

Sounds similar to the lawsuits against Google by the newspapers of the world for them providing excerpts of their articles on the search result page or so.

So essentially you can ask chatgpt to fetch and summarize a paywalled article, because openai has subscribed their crawler?

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#86
post #53

Earlier quoted context omitted.

I just tried asking my pencil to draw me a picture of Mario but nothing happened. What gives?

You're holding it wrong.

You need to use Google Images, and you will get millions of unauthorized electronic reproductions of Mario.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#88

So I can see the logic in treating the inputs to the AI training data sets the same way we treat humans learning something. A potential downside is that AI systems can 'mechanise' the creation of material that potentially infringes copyright (in the same way that human generated content can infringe) But a potential upside is that we can 'mechanise' the process by which we judge whether new content infringes the copy…

Generative AI tweaks the original material such that it's modified beyond recognition, and non verbatim. A bit like how artists tweak other artist's work, put their own spin on it, and then pass it off as 'original'.

If you ask it to sure, but there are examples of ChatGPT reproducing whole NYT articles, with not nearly enough alterations to constitute a new work.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#89

So I can see the logic in treating the inputs to the AI training data sets the same way we treat humans learning something. A potential downside is that AI systems can 'mechanise' the creation of material that potentially infringes copyright (in the same way that human generated content can infringe) But a potential upside is that we can 'mechanise' the process by which we judge whether new content infringes the copy…

Generative AI tweaks the original material such that it's modified beyond recognition, and non verbatim. A bit like how artists tweak other artist's work, put their own spin on it, and then pass it off as 'original'.

An alternative viewpoint:

Generative AI algorithmically processes the inputs such that they can be roughly recreated after decoding. A bit like how jpeg encodes images "beyond recognition"(aka lossy).

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#90
post #21

I think it’s a weak argument anyway. A pencil can be used to recreate copyrighted works. Why should a LLM be different? The legal responsibility always has been and should continue to be on the person publishing or making available the work. And at any rate, a country banning this tech will be missing the revolution. Protecting the buggy whip manufacturers and all that.

A pencil can be used to recreate copyrighted works. Why should a LLM be different? Speed and scale make these completely unrelated in my opinion.

[deleted]
Post reply on HN