Live data from Hacker News

Japan Goes All In: Copyright Doesn't Apply to AI Training

biia.com

141–150 of 183 posts

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#141
post #124
post #62

Earlier quoted context omitted.

My opinion is that people should just state that conclusion rather than make these silly comparisons.

Not everybody thinks in the same way. Some, maybe most people benefit from intuition pumps such as analogies. Good analogies clarify the point, poor ones obscure it.

>Good analogies clarify the point, poor ones obscure it.

I definitely agree with this, but I think a "good analogy" is really rare.

In this case, I think that the analogy just obscures the main point about copyright. I think the original comment would have been stronger by omitting the analogy and just discussing where the legal responsibility falls.

In almost every case an analogy is made, people discuss the validity of the analogy over the actual point of the comment (I'm guilty here too! my comment ended up being about the analogy rather than the main point).

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#142

Earlier quoted context omitted.

humans and AI aren't the same. There's no reason to think it will have to be the same.

They do have to be the same because learning is not something you can constrain with law pragmatically. Style and structure is not protected and can be mimicked so it is trivial to do parallel construction for any stylistic ideal even if the law says you can’t train on the original. Society does not benefit from these extra steps. Specific IP like characters are already perfectly well protected by copyright. It does…

You can train all the shit you want but if you train on someone else's materials without a license and your product makes money you should be forced to stop selling your product and to transfer the profits to the people who you took advantage of for your own I just enrichment. If you do it a lot you should go to jail.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#143

Earlier quoted context omitted.

humans and AI aren't the same. There's no reason to think it will have to be the same.

Yes there is - why should a well-proven law against plagiarism just evaporate if the human gets an ai to do it instead of does it themselves? Either the law should be revoked wholesale or it will have to be applied wholesale.

I don't think you understood what I wrote.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#144
post #64

Earlier quoted context omitted.

What's the difference between picking a cherry tomato from your streetside garden and eating it, and driving a combine harvester over your lawn and taking everything?

Pathetically incongruent analogy

Exactly, as was the previous one. Catching on to this stuff I see.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#145

Earlier quoted context omitted.

What's the difference between picking a cherry tomato from your streetside garden and eating it, and driving a combine harvester over your lawn and taking everything?

I'm not a lawmaker, but it's probably pretty hard to write a law that effectively distinguishes between "doing X" and "doing X at scale". As another commenter mentions, if you target the means of doing it (human doing X vs. machine doing X), someone will just use Mechanical Turk or something to hire 10,000 humans to do X. If telling AI to study Spiderman and then output 10 pictures of Spiderman is illegal, how is tha…

Not at all. We do it all the time. E.g. It's legal to consume fruits found in a National Park, as much as a single person can consume. It's illegal to harvest and take away anything more than that.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#146
post #41

There is a distinction to be made between inputs and outputs when it comes to AI and copyright. Most people focus on the former, and discuss whether you can train an LLM on copyrighted works or not, but ultimately the issues really only manifest in the latter. An AI training itself on a million newspaper articles can be declared legal, sure, but what happens when it also starts spitting out the same articles with nea…

> This is what the crux of NYT's lawsuit is about, and making laws about AI training isn't going to make a difference to that. The claims of NYT are more than just about training. They're claiming that as part of the ChatGPT software; it _looks up_ stuff in a database of articles. That is beyond _training_. If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think…

> it _looks up_ stuff in a database of articles.

That is clearly not the case. There isn't anything like enough storage space in typical LLMs to maintain a "database" of all the training data.

If an article can be reproduced from an extremely sparse representation using a stochastic algorithm, that seems to me to be prima facie evidence that the article didn't contain much (or possibly, any) significant creative content to begin with.

Certainly I would expect to find less creativity in an allegedly factual news article than in a piece of acknowledged fiction.

Only creative works are copyrightable in the United States.

There are soi-disant "artists" who produce "artworks" that are (e.g.) nothing but a pure white rectangle. That doesn't make any kind of copyright infringement.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#147
post #104
post #37

Earlier quoted context omitted.

Being able to train an AI is different than using it to reproduce copyright works verbatim. I could easily see a rule where you can train a model, but it's still on you to make sure that you don't use output from it that is too close to a copyright work. This to me seems like the right approach, and is not much different from what humans do. Humans are free to read whatever source material they want, but you can't su…

How can it possibly fall on you to verify your works? The only possible way to do this is for companies to provide a list of all of their sources and for me to then automate verification and hope it works! The real answer is that I should be able to control whether my information is used to train models or not, because we already know that models spit out verbatim results with generic queries and there’s just no way…

There is already have a blend of exactly what you desire in the US. But to answer your question, in the case of github copilot it does the checking for you, so no you don't need a list of all the sources.

Regarding you wanting to opt-out of training, that's fine you can already do that for many large models. But likely doing that will become the equivalent of putting your works in a safe where no one but you will ever end up reading them or finding them.

Banning all AI training is the equivalent to banning search with any modern search engine.

edit: Also, note that you are already on the hook for not infringing patents, which if you were to try to do "perfectly" like you imply with copyright, then you need to search/read/understand the entire body of published patents. A task that is clearly not possible. Yet patent law functions (unfortunately).

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#148
post #29

So I can see the logic in treating the inputs to the AI training data sets the same way we treat humans learning something. A potential downside is that AI systems can 'mechanise' the creation of material that potentially infringes copyright (in the same way that human generated content can infringe) But a potential upside is that we can 'mechanise' the process by which we judge whether new content infringes the copy…

> So I can see the logic in treating the inputs to the AI training data sets the same way we treat humans learning something. If we don't take this approach then there will be a series of very lame legal loop-holes with putting mechanical Turks [1] in the process. Or just end up with very "I know it when I see it" legislation. So for both practical and philosophical grounds I do support this. [1] https://en.wikipedia…

You don't even need that. Even "just" OpenAI has a valuation large enough to just buy several major publishers and data brokers to secure access to data if they need to license.

And while I expect NYT imagines that their archive is really valuable for training, they're just not that special in the sense that while they may have broken more stories on average than many others, and have had influential op eds etc., the ones that matters will have been cited and referenced and written about elsewhere - the irony is that by virtue of being so well known, their historically most important content is also less unique in terms of the accessibility of the information in it.

So while I'm sure OpenAI would love their archives, I'm also sure that if OpenAI and others have to license content and NYT end up being "difficult", OpenAI will just license content from (or buy) a suitably diverse portfolio of other papers instead.

In other words, beyond producing outright synthetic data, if AI companies are prevented from training on data they don't have a license to, the net effect will just be a scramble to buy licenses and/or buy companies that can provide sources of content, and the price for that content will be a lot lower than some of the people pursuing these copyright claims imagine.

In the end, if we go that route, all we'll have achieved as a society is creating massive moats protecting the companies already big enough to buy access to a broad enough set of content and made open models harder.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#149
post #84

Earlier quoted context omitted.

I would say it ought to be as legal as distributing a "how to draw mickey mouse" tutorial or a "how to sing taylor swift song" video or perhaps even a "how to make a twitter clone" tutorial.

In that case it seems like it would be much easier to just make it legal to distribute the copyright material in the first place.

That doesn't follow at all? Just because you can create something doesn't mean you can distribute it.

Just because I drew Mickey Mouse doesn't mean I can sell it. Just because I sing Taylor's song doesn't mean I can upload it to Spotify.

Just because GPT can return an image doesn't mean I'm allowed to sell it.

Creation and distribution are different, and the reality is that the 90% of consumers will never create, they will consume distribution.

Post reply on HN