Live data from Hacker News

Japan Goes All In: Copyright Doesn't Apply to AI Training

biia.com

131–140 of 183 posts

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#131

The source referenced is the minutes a representative made in a committee last year in april and within that a clarification question while discussing AI in education. It's not policy, not recent and not true.

But at least it was written by a real human! /s

But was it submitted by a real human? /s

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#132
post #102
post #61

Earlier quoted context omitted.

This is little more than baseless fearmongering. "zip up copyrighted images using a nn" is trek level technobabble. that's not how NNs work. how the hell does it even connect to privacy? copyright isn't what makes it illegal to expose and have your medical records it's privacy violations, which this doesn't even touch

> "zip up copyrighted images using a nn" is trek level technobabble. Look up 'overfitting', neural-network based compression, etc. or that paper that used zip compression as a neural-network basically. Farthest thing possible from being 'technobabble' once you understand how inextricably linked compression and 'understanding' is.

Regardless of whether it's "technobabble," it's a misunderstanding of how courts operate. The law is not a formally specified algorithm. If you overfit a NN to produce someone else's work, that's not going to get you off the hook in front of a court.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#133

Earlier quoted context omitted.

What's the difference between picking a cherry tomato from your streetside garden and eating it, and driving a combine harvester over your lawn and taking everything?

I'm not a lawmaker, but it's probably pretty hard to write a law that effectively distinguishes between "doing X" and "doing X at scale". As another commenter mentions, if you target the means of doing it (human doing X vs. machine doing X), someone will just use Mechanical Turk or something to hire 10,000 humans to do X. If telling AI to study Spiderman and then output 10 pictures of Spiderman is illegal, how is tha…

Isn't the latter already a copyright infringement? So in this case, both should be forbidden?

I think the more immediate issue with AI is that it's like having access to a close-to-zero cost human who doesn't care whether or not they're creating content which, if a human did it, would be considered a copyright infringement. And that they care so little about copyright (and other data rights) that they're basically incapable of even warning you if they are close to an existing character or living person.

I don't know how this is going to play out, but right now we're getting a more polite re-run of the luddites smashing early industrial equipment. I can sympathise with the loss of purpose and economic disenfranchisement, but the economic power in that revolution went to those that did the most automation, and I expect the same to be true this time.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#134
post #85

Earlier quoted context omitted.

> This is what the crux of NYT's lawsuit is about, and making laws about AI training isn't going to make a difference to that. The claims of NYT are more than just about training. They're claiming that as part of the ChatGPT software; it _looks up_ stuff in a database of articles. That is beyond _training_. If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think…

Sounds similar to the lawsuits against Google by the newspapers of the world for them providing excerpts of their articles on the search result page or so. So essentially you can ask chatgpt to fetch and summarize a paywalled article, because openai has subscribed their crawler?

Which lawsuits? Newspapers are free to remove themselves from google at any time.

You might be thinking of newspapers’ political campaigns such as C-18 in Canada, which were done in bad faith. The newspapers wanted links to stay but wanted money too.

OpenAI can’t do this for NYT without destroying their model and remaking it.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#135

For all we know, in 5 years you could ask an AI to "reproduce the entire Supernatural TV show but change the character names, dialog and look just enough to bypass copyright issues - 7 seasons, 24 episodes each MP4 format"

Seven seasons? I thought you said reproduce the entire show? 37 seasons :D

I see we have a fan of that show! :)

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#136
post #41

There is a distinction to be made between inputs and outputs when it comes to AI and copyright. Most people focus on the former, and discuss whether you can train an LLM on copyrighted works or not, but ultimately the issues really only manifest in the latter. An AI training itself on a million newspaper articles can be declared legal, sure, but what happens when it also starts spitting out the same articles with nea…

> An AI training itself on a million newspaper articles can be declared legal, sure, but what happens when it also starts spitting out the same articles with nearly no modification? Is that still fair use?

That would not be transformative. If it is also provided as a commercial service that directly competes against the copyright holder, then that would not be 'fair use'.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#137

Earlier quoted context omitted.

> This is what the crux of NYT's lawsuit is about, and making laws about AI training isn't going to make a difference to that. The claims of NYT are more than just about training. They're claiming that as part of the ChatGPT software; it _looks up_ stuff in a database of articles. That is beyond _training_. If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think…

> If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think everybody would agree that's copyright infringement. Sure. Slightly more interesting is if that same human with those same breifcases was taking money to answer questions and referenced those papers, but did not just provide the article or headlines, and might not even be paraphrasing the article at all. I…

While LLMs are designed to generalize from their training data, rather than simply memorizing it, overfitting occurs in niche areas. There's be plenty of niche areas and LLMs of today merely repeat training data, like with the NYT case. Larger datasets and better algorithms will help to an extent, but you'll always have niche topics where overfitting happens. The intent has always been to generalize well, however, it may not be feasible to do so in the long tail of the internet. How should copyright law address this?

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#138

So I can see the logic in treating the inputs to the AI training data sets the same way we treat humans learning something. A potential downside is that AI systems can 'mechanise' the creation of material that potentially infringes copyright (in the same way that human generated content can infringe) But a potential upside is that we can 'mechanise' the process by which we judge whether new content infringes the copy…

Generative AI tweaks the original material such that it's modified beyond recognition, and non verbatim. A bit like how artists tweak other artist's work, put their own spin on it, and then pass it off as 'original'.

GenAI CAN do that, but it also can regenerate training sources as-is. The way LLMs work makes it quite likely that once it has started copying a training sample it will continue to do so (after copying/generating N consecutive words of a training sample, the most statistically likely next word will be often be the N+1th word of that same sample).

The same thing happens for images - the more an output looks like an input, the more likely (as diffusion-based generation proceeds) it is to look more like it. There are recent examples of Dune posters being recreated essentially as-is.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#139

The source referenced is the minutes a representative made in a committee last year in april and within that a clarification question while discussing AI in education. It's not policy, not recent and not true.

Plus this website looks so SEO'ed that it's unclear whether it's even a real 'association'.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#140

Earlier quoted context omitted.

Generative AI tweaks the original material such that it's modified beyond recognition, and non verbatim. A bit like how artists tweak other artist's work, put their own spin on it, and then pass it off as 'original'.

If you ask it to sure, but there are examples of ChatGPT reproducing whole NYT articles, with not nearly enough alterations to constitute a new work.

What were the prompts? A whole article would barely fit in the output tokens, typically, yeah?
Post reply on HN