Live data from Hacker News

AI is in danger of being swallowed up by copyright law

heathermeeker.com

451–460 of 705 posts

Re: AI is in danger of being swallowed up by copyright law

#451

Earlier quoted context omitted.

It's not plagiarism at all. The AI is trained on 5 billion images yet it stores only 4gb of data. Thus it is impossible that it stores the actual work. For any image that the AI generates, you can't point to any image in the training data that the image is derived from.

How did they train the AI without first storing the data? It's not in the model, but it was used without permission in the pipeline that lead to creating that AI model. I don't know if that counts as plagiarism, but there's clearly some use of this copyright material that the authors probably didn't envision and did not grant permission for. I have no idea what the law would be in cases like this

> How did they train the AI without first storing the data?

the data was originally permitted to be copied.

The question isn't whether the training is violating copyright - as long as the data set had permission to be viewed (which it must have, since it was public).

The question is whether the final result - the model/weights - is a derivative work of the training data set. If it is a derivative work, then the model must be in violation of copyright. But copyright law allows for sufficiently transformative work to be considered new, rather than derivative. So is training a model using methods like this constitute a transformative work?

Re: AI is in danger of being swallowed up by copyright law

#452

Earlier quoted context omitted.

> Do you ask for permission when you train your mind on copyrighted books yes, thats why I pay a fee to buy/borrow one (or someone pays the fee in the case of a library.) > listen to music again money is exchanged. > Humans are constantly ingesting gobs of "copyrighted" insights that they eventually remix into their own creations without necessarily reimbursing the original source(s) of their creativity. yes, and so…

You don't need to pay for or "borrow" anything to learn from copyrighted works. Nobody has had that expectation for years, and that is also not what copyright pertains to. It's not that AI breaks into libraries and isn't paying the fees. You can google an image of any great work of art and look at it for as long as you like, for free and take from it what you can and use all of that to create something else and get p…

> You don't need to pay for or "borrow" anything to learn from copyrighted works.

someone pays, just maybe not you. How do you think google/meta/et al offer you a service free at the point of delivery, through charity?

> You can google an image of any great work of art and look at it for as long as you like

see my bit about google. The copyright still is with the owner. That image can be removed, should the owner wish, but for various reasons its too expensive to get google to respect that.

> I would argue most people are quite happy with the existence of something like Google search

yes, because its a symbiotic relationship. I as a creator, make something that people want to find, google points them to me, and I get people's attention. I might do that to fluff my ego, or try and convert it to cash through sales or something.

The AI step threatens to remove that relationship. Instead of being passed to me, the AI just pastes shit its gleaned from mine and other websites, leaving no chance of me getting a reward for making that website.

Re: AI is in danger of being swallowed up by copyright law

#454
post #121

Earlier quoted context omitted.

AI companies can ask for permission if they want to train their models on other people's works Do you ask for permission when you train your mind on copyrighted books? Or observe paintings? Or listen to music? Do you ask for permission when you get new ideas from HN that aren't your own? Humans are constantly ingesting gobs of "copyrighted" insights that they eventually remix into their own creations without necessar…

> Do you ask for permission when you train your mind on copyrighted books? The law already makes many distinctions between humans and machines. For example, looking out the window to see when your neighbor is going to the supermarket: allowed; using a machine-vision system to store the movements of groups of people into a large database: not allowed. Also, "training the mind" and "training a machine learning system"…

This is the crux of the issue, I believe.

It seems to me that one side is arguing that people (as in, individual human beings) already do what the AI is being accused of, the other side argues that it's replicating work.

The truth of the matter is that what is taking place is a different thing altogether. We do generally deal in a different way with "machine behavior" because we recognize it being automatic and reproducible matters.

Re: AI is in danger of being swallowed up by copyright law

#455
post #418

There's no part of AI that is being swallowed up by copyright. AI companies can ask for permission if they want to train their models on other people's works. It's not that hard, various image hosting sites have already added an opt-in/opt-out toggle to their services. Sites might even get away with using this stuff as compensation for free hosting. The fact of the matter is that the AI companies don't want to ask fo…

You don’t need permission. If you want to prove your data was used to train an AI, the onus is on you to prove it. Good luck. The AI who follow the law strictly will be at a disadvantage to those that do not.

> the onus is on you to prove it

which would be easy during a law suit - the process of discovery means you get to check out the training dataset.

The allegation isn't that the AI trainers are hiding, but that what AI trainers are doing _itself_ constitutes copyright violation. AKA, they want the right to use the works to train an ai model to be a right that must be explicitly granted.

Re: AI is in danger of being swallowed up by copyright law

#456
post #238

US copyright law only applies in the US. AI can be developed outside of the US where other laws apply. So, the US using the straight jacket of copyright law to stifle innovation would only cause AI companies to move their business; or non US companies stepping up. I don't think that's actually going to happen though. The stakes are too high for that and this case seems quite weak. Copyright law has its limitations. B…

Where? Even Afghanistan’s joined the WTO. Wouldn’t the US file a WTO dispute if that’s the way the wind blows? The harm to US industries would be enormous.

Belarus. Haven't you seen the news of this month?

Re: AI is in danger of being swallowed up by copyright law

#457

Earlier quoted context omitted.

I'm sure chatgpt would be of much less interest worldwide if it only spoke japanese.

Doesn't follow. Why would a tool built in Japan limit itself to Japanese? ChatGPT doesn't limit itself to English.

To export you must abide to the laws of the country you're exporting to.

Re: AI is in danger of being swallowed up by copyright law

#458

Earlier quoted context omitted.

If it really turns out to be a problem then copyright holders are simply going to add in a not-licensed-for-training clause in all their licenses[1]. Sure, existing works already licensed can still be used, but at least both parties (copyright holders and AI trainers) won't have anything to argue about. [1] Anyone from CC reading this? Make it the default.

It is like adding a clause in the license that you are not allowed to read the license. The moment you share your creation/work to someone/the world, you are training their nn. You can not share something publicly and then demand "xyz" can not view it. Viewing is training. You are free to keep your creation under lock & key and only share with nn (of people and/or AI) of your choice.

> You can not share something publicly and then demand "xyz" can not view it.

That's nonsense. Licenses have clauses on how the content may be used. Clauses along the lines of "The content may not be used for ..." are common.

I dunno where you heard that once you release something the license clauses no longer apply, but it's wrong.

Re: AI is in danger of being swallowed up by copyright law

#459

Earlier quoted context omitted.

> Do you ask for permission when you train your mind on copyrighted books? Or observe paintings? Or listen to music? Yes, that’s exactly what happens when you buy a book, or pay for a music subscription. The work is in the public domain, then global permission to observe and copy the work is already granted. > Do you ask for permission when you get new ideas from HN that aren't your own? You don’t need to. It’s impli…

>Yes, that’s exactly what happens when you buy a book, or pay for a music subscription. The work is in the public domain, then global permission to observe and copy the work is already granted. When you buy a book, you’re not paying a licensing fee. You’re exchanging for goods. You’re granted very few rights to own a copy of the work. But they’re almost all to do with distribution. None of those rights is the right t…

When you buy a book or some other artwork, it is implicitly assumed you will put it in your brain, or your meat neural net. And that your brain could produce something related to this content.

It's not just assumed, it's celebrated when a work of art gathers fans who produce their own, inspired content.

Not sure why it needs to be over-complicated or different for silicone neural nets. But I think it will get very over-complicated, if not politicised, in the following years.

Re: AI is in danger of being swallowed up by copyright law

#460
post #79

Earlier quoted context omitted.

As an extension of this, only allow children to look at works they purchased publication rights to, lest their creative output becomes influenced by a different person's style.

I have no idea what your point is here. These AI companies are making serious amounts of money (OpenAI is valued in tens of billions) on the back of artists who never gave permission for their work to be used in this way. If a child took an artist's work, copied it and made significant amounts of money from selling it then yes they should be within the purview of copyright law.

> who never gave permission for their work to be used in this way.

the copyright aren't all encompassing. There's only an enumerated set of rights granted, and "this way" (aka, training an AI model) is not one of those restricted activities (like distribution or broadcast).

Unless the model can be argued to be a derivative work of the training data set (which i don't believe it is, since the process of training is sufficiently transformative imho), the original copyright holders of the training data do not need to be asked permission.

Post reply on HN