Live data from Hacker News

Japan Goes All In: Copyright Doesn't Apply to AI Training

biia.com

91–100 of 183 posts

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#92

The source referenced is the minutes a representative made in a committee last year in april and within that a clarification question while discussing AI in education. It's not policy, not recent and not true.

That's how I like all of my news stories - untrue, untimely, and based off of the word of one person in the legislature.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#93
post #76

I think it's notable that when the copyright holder is a business there is rarely any questioning that the copyright holder is king. But when the copyright holder is an individual or an artist, then it practically becomes a challenge for the government and/or private sector to try and strip that copyright away from you.

That’s because of publishing, mostly. Individual creators very often sign over some or all control of their work to companies in order to make more money.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#94
post #62
post #36

Earlier quoted context omitted.

Speed and scale also makes the printing press completely different than the pencil, but the principle that the user is responsible for the output remains practical.

My opinion is that people should just state that conclusion rather than make these silly comparisons.

[deleted]

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#96
post #29

So I can see the logic in treating the inputs to the AI training data sets the same way we treat humans learning something. A potential downside is that AI systems can 'mechanise' the creation of material that potentially infringes copyright (in the same way that human generated content can infringe) But a potential upside is that we can 'mechanise' the process by which we judge whether new content infringes the copy…

> So I can see the logic in treating the inputs to the AI training data sets the same way we treat humans learning something. If we don't take this approach then there will be a series of very lame legal loop-holes with putting mechanical Turks [1] in the process. Or just end up with very "I know it when I see it" legislation. So for both practical and philosophical grounds I do support this. [1] https://en.wikipedia…

And "clean room" AI training other AI.

The end is inevitable, protecting copyright for AI training is probably a lost cause.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#97
post #85

Earlier quoted context omitted.

> This is what the crux of NYT's lawsuit is about, and making laws about AI training isn't going to make a difference to that. The claims of NYT are more than just about training. They're claiming that as part of the ChatGPT software; it _looks up_ stuff in a database of articles. That is beyond _training_. If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think…

Sounds similar to the lawsuits against Google by the newspapers of the world for them providing excerpts of their articles on the search result page or so. So essentially you can ask chatgpt to fetch and summarize a paywalled article, because openai has subscribed their crawler?

[deleted]

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#98
post #68
post #57

Earlier quoted context omitted.

So if I use 500 pages/minute scanner it will be fine?

My comment doesn't say what my opinion on the topic is, just that pencils and LLMs are hardly comparable due to difference in speed and scale.

Congratulations, by attaching qualitative legal importance to scale, speed, and size, you just successfully argued that the First Amendment doesn't apply on the Internet.

You'll make friends on both sides of the political aisle with that position, but it's a shame to see it so readily accepted around here.

(Can't reply due to HN's rate limiting algorithm that penalizes me after four or five posts while allowing the most-corrosive trolls imaginable to party all day, so I edited to clarify.)

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#99
post #41

There is a distinction to be made between inputs and outputs when it comes to AI and copyright. Most people focus on the former, and discuss whether you can train an LLM on copyrighted works or not, but ultimately the issues really only manifest in the latter. An AI training itself on a million newspaper articles can be declared legal, sure, but what happens when it also starts spitting out the same articles with nea…

> This is what the crux of NYT's lawsuit is about, and making laws about AI training isn't going to make a difference to that. The claims of NYT are more than just about training. They're claiming that as part of the ChatGPT software; it _looks up_ stuff in a database of articles. That is beyond _training_. If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think…

> If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think everybody would agree that's copyright infringement.

Except that's not what it's doing. Show me how I can get ChatGPT to show me the full text of a NYT article.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#100
post #77
post #46

Earlier quoted context omitted.

It'd be really interesting to open up a movie theater in Japan that just ingests blockbusters through a "do nothing" NN and then be able to screen them royalty free. This decision feels incredibly half baked.

This is explicitly for training, not for the distribution of copyrighted material. Courts aren't stupid.

I'd clarify that in my example the screening would be of the potentially random output of a model... just one that was only trained by watching a specific blockbuster movie and thus extremely likely to just reproduce the source material. My example is obviously an extreme but it gets at the core of the NYT case here in the states... I think it's a bad thing if we allow models to output data nearly indistinguishable from copyrighted data it was trained on.

W.r.t to the NYT case - It's my opinion that it's completely reasonable to use a corpus of vetted english literature like the NYT as a way to train your model to comprehend language - but if the model also begins to echo the contents of those articles then that may be a serious breech of the NYT's right's to monetize their work.

Post reply on HN