Live data from Hacker News

Japan Goes All In: Copyright Doesn't Apply to AI Training

biia.com

121–130 of 183 posts

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#121

The source referenced is the minutes a representative made in a committee last year in april and within that a clarification question while discussing AI in education. It's not policy, not recent and not true.

But at least it was written by a real human! /s

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#122

So I can see the logic in treating the inputs to the AI training data sets the same way we treat humans learning something. A potential downside is that AI systems can 'mechanise' the creation of material that potentially infringes copyright (in the same way that human generated content can infringe) But a potential upside is that we can 'mechanise' the process by which we judge whether new content infringes the copy…

The issue is the output, not the input. If LLMs were just learning from their training set and generating novel output, the same way a human might, then there'd be no problem.

The issue is that generative AI - both for images as well as text - in effect memorize sources as well as learn from from, and can end up regenerating training sources verbatim (or with minimal changes in case of images).

I don't think any US court is going to accept "yes your honor, we copied this copyright material, but we used a TOOL to do it" as a way to avoid copyright.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#123
post #41

There is a distinction to be made between inputs and outputs when it comes to AI and copyright. Most people focus on the former, and discuss whether you can train an LLM on copyrighted works or not, but ultimately the issues really only manifest in the latter. An AI training itself on a million newspaper articles can be declared legal, sure, but what happens when it also starts spitting out the same articles with nea…

> nearly no modification How much modification is enough modification? The courts can wrangle with that endlessly.

Sure, and that's the point of having courts. If you look at the examples in NYT's filing though I don't think anyone can argue that it isn't clear-cut plagiarism.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#124
post #62
post #36

Earlier quoted context omitted.

Speed and scale also makes the printing press completely different than the pencil, but the principle that the user is responsible for the output remains practical.

My opinion is that people should just state that conclusion rather than make these silly comparisons.

Not everybody thinks in the same way. Some, maybe most people benefit from intuition pumps such as analogies. Good analogies clarify the point, poor ones obscure it.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#125

Earlier quoted context omitted.

Completely crazy take. Your health records are not currently protected by copyright law. If someone magically snapped their fingers and eliminated copyright law, the protections on your health data, scant though they may be, would more or less be the same (IANAL)

So to be clear -- you're well aware that in the case of your private data that you have an interest in preventing it being used to train AI. Great. So d'you think you could outline a reason why you wouldnt have an interest in your creative works not being used also? Either the training data is, as big-ad-tech says, essentially equivalent to generic human experiences -- ie., weakly repoducible; OR it is extremely repo…

Because my private data isn't protected by copyright, it's protected by things like HIPAA which doesn't matter one iota about human experiences and applies equally to humans and machines. It's about data sovereignty and who may access my data and for what purpose. A human is not allowed to share, retain, or reproduce my medical data.

So arguments like "I can get the AI to output my chart verbatim" start carrying weight because it's granted access to data that the humans that created the AI are not permitted to share in any form whatsoever where as copyright concerns what I may do with the data after it's produced. Copyright is full of exceptions for things that don't count as a reproduction or performance of the work and this is just one more, it doesn't change the nature of copyright.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#126
post #85

Earlier quoted context omitted.

> This is what the crux of NYT's lawsuit is about, and making laws about AI training isn't going to make a difference to that. The claims of NYT are more than just about training. They're claiming that as part of the ChatGPT software; it _looks up_ stuff in a database of articles. That is beyond _training_. If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think…

Sounds similar to the lawsuits against Google by the newspapers of the world for them providing excerpts of their articles on the search result page or so. So essentially you can ask chatgpt to fetch and summarize a paywalled article, because openai has subscribed their crawler?

I believe it really only works with GPT-4 directly, because OpenAI's prompting to make ChatGPT a chatbot ruins the effect.

Specifically the NYT put in the first sentence of the article and asked GPT-4 to autocomplete it, which it did with >95% accuracy. It's really quite stark: https://nitter.net/jason_kint/status/1740146134767865895#m

The issue isn't that GPT-4 users can read NYT stories without paying for a subscription, though that is a legitimate concern. The issue is that for good-faith use cases - e.g. asking to write a summary about a recent current event - GPT-4 could very well copy an entire paragraph from NYT verbatim, without the user having any way of knowing. It's a serious problem.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#127
post #115
post #42

Earlier quoted context omitted.

It seems pretty clear what AI means in this context. > There is no difference between a lossy jpg, taking its pixels as weights, and the weights of a NN. Somehow I don't think this is going to hold up in court.

At that direct level yeah probably, and I do think copyright is dumb and should be if not abolished, limited to 20 years or whatever. That said, imagine training a network from scratch entirely on Disney's catelog. Even if that model is then prompted to generate new characters, it seems weird to say that Disney's copyright wasn't infringed.

[deleted]

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#128
post #85

Earlier quoted context omitted.

> This is what the crux of NYT's lawsuit is about, and making laws about AI training isn't going to make a difference to that. The claims of NYT are more than just about training. They're claiming that as part of the ChatGPT software; it _looks up_ stuff in a database of articles. That is beyond _training_. If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think…

Sounds similar to the lawsuits against Google by the newspapers of the world for them providing excerpts of their articles on the search result page or so. So essentially you can ask chatgpt to fetch and summarize a paywalled article, because openai has subscribed their crawler?

[deleted]

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#129

Earlier quoted context omitted.

> This is what the crux of NYT's lawsuit is about, and making laws about AI training isn't going to make a difference to that. The claims of NYT are more than just about training. They're claiming that as part of the ChatGPT software; it _looks up_ stuff in a database of articles. That is beyond _training_. If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think…

> If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think everybody would agree that's copyright infringement. Except that's not what it's doing. Show me how I can get ChatGPT to show me the full text of a NYT article.

> Show me how I can get ChatGPT to show me the full text of a NYT article.

There are like 40 pages of examples in NYT's lawsuit showing exactly that.

Post reply on HN