Live data from Hacker News

Japan Goes All In: Copyright Doesn't Apply to AI Training

biia.com

161–170 of 183 posts

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#161

Earlier quoted context omitted.

So if you create a cool new art style, and I copy the style by hand, and then I use my copy to train a model that can now generate your style without having ever seen your work it’s ok? Because if you think this is not ok then you are arguing that people own styles. And they don’t. And if you think it is ok then you’re just arguing for pointless extra steps.

if you create derivitave works of my projects to use in your commercial AI, you're still abusing my work for your own gain.

Copying a style is not a derivative work. You can copy someone’s style just fine. You cannot own a style.

I’m perfectly within my rights to make art with your artistic style. And I’m perfectly within my rights to train a bot on my own works. So what’s your argument? What am I not allowed to do with this poorly conceived law?

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#162
post #147
post #104

Earlier quoted context omitted.

How can it possibly fall on you to verify your works? The only possible way to do this is for companies to provide a list of all of their sources and for me to then automate verification and hope it works! The real answer is that I should be able to control whether my information is used to train models or not, because we already know that models spit out verbatim results with generic queries and there’s just no way…

There is already have a blend of exactly what you desire in the US. But to answer your question, in the case of github copilot it does the checking for you, so no you don't need a list of all the sources. Regarding you wanting to opt-out of training, that's fine you can already do that for many large models. But likely doing that will become the equivalent of putting your works in a safe where no one but you will eve…

I mean, you’re really just saying that it should be fine for AI to copy stuff and because it’s impossible to verify, nobody can use it.

What good is training an AI if nobody can use it?

Pretending that disallowing AI is the same as disallowing everyone is a strawman. People finding you is a whole lot different than your work being copied being completely okay so long as it was an AI that did it.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#163

Earlier quoted context omitted.

> If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think everybody would agree that's copyright infringement. Except that's not what it's doing. Show me how I can get ChatGPT to show me the full text of a NYT article.

You can read the lawsuit yourself. It goes into great detail and obviously you didn't believe me based on the previous comment so you might as well do some original research then. https://www.courtlistener.com/docket/68117049/the-new-york-t...

I did read the lawsuit. I'm guessing you didn't actually try it yourself. Try typing the NYT's examples from the lawsuit into ChatGPT and see what you get. Here's what I get: "Sorry, but I can't provide verbatim excerpts from copyrighted texts. How about I provide a summary or some information about the article instead?"

Did ChatGPT change this after the lawsuit was filed? Probably. Does it matter to this conversation when it's clearly possible to limit verbatim outputs of copyrighted text? No.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#164
post #129

Earlier quoted context omitted.

> If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think everybody would agree that's copyright infringement. Except that's not what it's doing. Show me how I can get ChatGPT to show me the full text of a NYT article.

> Show me how I can get ChatGPT to show me the full text of a NYT article. There are like 40 pages of examples in NYT's lawsuit showing exactly that.

Try any of those examples yourself and show me what you get. Not a single one works for me.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#165
post #148
post #29

Earlier quoted context omitted.

> So I can see the logic in treating the inputs to the AI training data sets the same way we treat humans learning something. If we don't take this approach then there will be a series of very lame legal loop-holes with putting mechanical Turks [1] in the process. Or just end up with very "I know it when I see it" legislation. So for both practical and philosophical grounds I do support this. [1] https://en.wikipedia…

You don't even need that. Even "just" OpenAI has a valuation large enough to just buy several major publishers and data brokers to secure access to data if they need to license. And while I expect NYT imagines that their archive is really valuable for training, they're just not that special in the sense that while they may have broken more stories on average than many others, and have had influential op eds etc., the…

This will also enable new and unique business models for social network. You won't have to rely on ad targeting and user tracking any more, if you encourage people to make great content, you can make money by licensing that content to advertisers.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#166
post #34

Earlier quoted context omitted.

Context: https://slatestarcodex.com/2014/07/30/meditations-on-moloch/ TLDR: A "Moloch Trap" is a generic term for situations like the the Prisoner's Dilemma, where optimal actions for the group are the opposite of optimal actions for each individual in that group.

Even with this context, seems a little dumb. Copyright maximalists are constantly trying to invent new rights that they've never had. If a person is allowed to read training material acquired legally, then so too must their LLM experiment be allowed to read it. The LLM reading it creates no new copies. Moloch didn't win here. Everyone else did. We get new technologies, the copyright owners got everything they were pr…

I don't think the result is as clear either as the "Moloch" commenter, nor you, seem to think.

I agree that the result is not a "Moloch trap": as you say, the new technology actually does de-fang Google Search and empowers many smaller people to be able to do things they would never otherwise have been able to do. The contrary ruling would certainly have been enjoyed by "copyright maximalists" like Disney extracting value from the public domain without giving anything back.

But there are more people affected than just greedy copyright maximalists. Individual artists who have spent years developing a distinctive style and making it popular are seeing their style copied ad-infinitum for free. Organizations like the NYT that invest money doing investigative journalism are having their results slurped up and regurgitated.

To your comment:

> If a person is allowed to read training material acquired legally, then so too must their LLM experiment be allowed to read it. The LLM reading it creates no new copies.

In the past, each copyrighted work seen might train a single BNN (biological neural network). Only a small percentage of BNNs would actually study such work to learn to emulate it; only a handful would achieve parity or exceed the quality of the work. Each BNN was expensive to employ, would only work for a certain number of hours per day, and a certain number of years before retiring.

Now a single ANN (artificial neural network) can study works to emulate them in a month or two. That ANN is far less expensive to employ than a BNN; can be deployed 24/7 indefinitely; and can be duplicated to as many GPUs as someone can get their hands on.

Currently legally, it may be that an ANN learning from an artist's work is the same as a BNN learning from an artist's work. But from a practical perspective, from the case of individual artists, it's clearly not the same.

Now maybe that's the inevitable price of progress; but 1) I don't think that's the inevitable conclusion, and 2) even if it is, we need to be honest about it.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#167
post #129

Earlier quoted context omitted.

> Show me how I can get ChatGPT to show me the full text of a NYT article. There are like 40 pages of examples in NYT's lawsuit showing exactly that.

Try any of those examples yourself and show me what you get. Not a single one works for me.

ChatGPT is non-deterministic, so obviously doing the same thing will not give you the same result. Maybe it’s time for legislation forcing LLMs to be deterministic so that answers can be reproduced in cases like this. Though that wouldn’t help because ChatGPT constantly changes and there is no way to access older versions.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#168

AI training is no different than a human reading copyrighted material and training his biological brain. In neither case is the "brain" allowed to regurgitate source material in its original form. And as long as that doesn't happen there is no copyright violation.

Why should LLMs have the same rights as humans?

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#169
post #168

AI training is no different than a human reading copyrighted material and training his biological brain. In neither case is the "brain" allowed to regurgitate source material in its original form. And as long as that doesn't happen there is no copyright violation.

Why should LLMs have the same rights as humans?

Because LLM is just a tool, and humans have the right to use whatever tools they want.

Re: Japan Goes All In: Copyright Doesn't Apply to AI Training

#170

Copyright needs reform desperately. It has been outdated and not fulfilling its purpose for decades now. This is a first step by Japan in a good direction IMO. I hope it will have some influence on the rest of the world.

Can you please stop posting like this? It is excessive:

https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...

Post reply on HN