Live data from Hacker News

AI is in danger of being swallowed up by copyright law

heathermeeker.com

461–470 of 705 posts

Re: AI is in danger of being swallowed up by copyright law

#461
post #443

Earlier quoted context omitted.

> Just because you can view something or free doesn't mean you can use it anyway you want. This whole thread really makes me want to pull my hair out. Difference between illegaly creating a (even temporary) copy of a copyrighted work (e.g. streaming a movie) vs. creating a derivative work of said copyrighted work: Two completely different things, with completely different legal outcomes. If OpenAI in any shape or for…

> anything that a computer does can be construed as making a copy. but that's not the point of contention. The training data set has been granted the right to be distributed (by virtue of it being available for viewing already - it's not hidden or secret). The proof is that a human can already view it manually. Let's call this 'public'. The question is, whether using this public training dataset constitutes creating…

>but that's not the point of contention. The training data set has been granted the right to be distributed (by virtue of it being available for viewing already - it's not hidden or secret). The proof is that a human can already view it manually. Let's call this 'public'.

This is wrong. My paintings are publicly available (especially going by your definition [which I'm confused by the origin of?]). Taking a photograph of my paintings is still a copyright violation. I hope we can ignore all the legal kerfuffle about personal use, as it has no bearing on our discussion. Again -- all of this boils down back to what I've said before: Bare human consumption does not constitute as making a copy, nearly everything else does.

Your second point -- a copyrighted work automatically granting someone else any rights (especially distrubtionial rights) by just being available to be consumed -- is even more wrong. I'm not going to go further into that, as you can very easily prove yourself wrong by googling it.

>The question is, whether using this public training dataset constitutes creating a derivative work

I'm not well versed in the US copyright laws, but I would assume (strongly so) that this would not be the case. I -- again, for US copyright law -- assume that for something to be considered a derivative work, it needs to include (or be present in other ways) copyrightable (!) parts of the original work(s). In other words, the original work needs to "shine through" the derivative work, in one way or the other. The delta of parameter changes of a ML model would (imo) not constitute such a thing.

Problems with derivative works will come into play when considering the things ML models produce.

Re: AI is in danger of being swallowed up by copyright law

#462

Earlier quoted context omitted.

Regardless of whether one agrees or not with paying creators of the training data, I think the deeper issue here is about societal wealth distribution and who gets paid for X now that X is being done very well by AIs. A less equitable world has Google or billionaires getting paid. A more equitable world has the artists. But I want to argue here that for purposes of this latter question, your proposal of copyright enf…

> - Even if you get the system to work, what about future artists and writers? Are we just creating an entrenched historical group of creatives getting royalties forever? The boat has long since sailed on this… ands it’s globally entrenched as a norm of international trade that we are all “ok with this” regime of 75 years or century plus copyright terms … And arguably the entire copyright vs AI/ML training datasets d…

What’s great about your comment is that you show that what we need the most is to reduce the power of copyright in such a corporate-centric legal regime.

Personally I’d like to see the right to train statistical models on any works without the permission of the author enshrined in statute and an end to common-law copyright, a return to the Statute of Anne 14/28 time length, and a clear delineation between the “work” as having an author for an eternity but having a “copyright of the work” vastly limited in scope.

Ask yourself, do we want to be extending the reach of large copyright holders like Disney into taking a fee from LLM producers because they COULD be helping people draw Mickey ears on their private creations?

This is Betamax all over again and luckily that Supreme Court opinion will favor heavily in the lower court’s judgement of these models as fair use.

Re: AI is in danger of being swallowed up by copyright law

#463
post #88
post #83

Earlier quoted context omitted.

I can instruct an AI to draw Mickey Mouse and infringe copyright and I can instruct a child to draw Mickey Mouse and infringe copyright.

The problem is there's no way a user will know whose copyright they are infringing when they ask AI to "paint a landscape." Maybe the AI needs to be able to print out a list of sources to provide attribution. That would be interesting.

> they ask AI to "paint a landscape."

that's the responsibility of the user said AI to check.

Re: AI is in danger of being swallowed up by copyright law

#464
post #121

Earlier quoted context omitted.

AI companies can ask for permission if they want to train their models on other people's works Do you ask for permission when you train your mind on copyrighted books? Or observe paintings? Or listen to music? Do you ask for permission when you get new ideas from HN that aren't your own? Humans are constantly ingesting gobs of "copyrighted" insights that they eventually remix into their own creations without necessar…

> Do you ask for permission when you train your mind on copyrighted books yes, thats why I pay a fee to buy/borrow one (or someone pays the fee in the case of a library.) > listen to music again money is exchanged. > Humans are constantly ingesting gobs of "copyrighted" insights that they eventually remix into their own creations without necessarily reimbursing the original source(s) of their creativity. yes, and so…

>Culture isn't free. Someone is paying for it, and if you stop paying them, then it doesn't get created.

This is just deeply wrong. Culture existed before money. It is tragic to me that a person can't see culture as anything but a marketable good.

Re: AI is in danger of being swallowed up by copyright law

#465

Earlier quoted context omitted.

If it really turns out to be a problem then copyright holders are simply going to add in a not-licensed-for-training clause in all their licenses[1]. Sure, existing works already licensed can still be used, but at least both parties (copyright holders and AI trainers) won't have anything to argue about. [1] Anyone from CC reading this? Make it the default.

How are you planning on proving a particular licensed work was used in a sufficiently large model? One of the commonly mentioned issues with current ML is the inability to reverse the output to figure out 'how it got there'.

> How are you planning on proving a particular licensed work was used in a sufficiently large model? One of the commonly mentioned issues with current ML is the inability to reverse the output to figure out 'how it got there'.

That's a different problem. Let's not get into the argument of "Just because the victim cannot prove something, we should remove the relevant laws."

The current laws are sufficient; all that it takes is for licenses to have a non-AI-training clause.

Lets solve the problem of "how do you prove" when we get to it[1].

[1] Right now, due to the systems already trained being given every single image on the net as training data according to the owner of those systems, in a civil suit the burden will be on them to prove that, on the balance of probabilities, a particular image was not used.

Re: AI is in danger of being swallowed up by copyright law

#466
post #249

Earlier quoted context omitted.

All or nothing, in my opinion. Either abolish or severely reduce copyright, or abide by it. The simple fact of the matter is that Disney and Getty invest a lot of money into these materials being out there in the first place. Open source programmers and artists spend a lot of time producing works for no cost other than some minor courtesies. AI companies aren't your friend or the little mom ''n pop shop down the road…

> All or nothing, in my opinion. Either abolish or severely reduce copyright, or abide by it. I firmly believe in "practice what you preach". I you declare you firmly believe in A but then do something directly counter to that because it's more convenient in this specific case, then that doesn't sit right with me. Besides, further expanding copyright in this one area will only make it so much harder to reduce it late…

that's like saying "you say you don't believe in borders, yet you oppose this invasion? curious"

Re: AI is in danger of being swallowed up by copyright law

#467

Earlier quoted context omitted.

Yes but using AI to generate works that can be used commercially is the way commercial AI companies plan to monetize AI.

There have been leaks where OpenAI is charging $42/month to use their service. How much of that is going back to the copyright holders whose work their service derives value from ?

> How much of that is going back to the copyright holders whose work their service derives value from ?

how much of the earnings of the student of art goes to the textbook authors, paintings and learning materials he used to get to where he is today?

Re: AI is in danger of being swallowed up by copyright law

#468

Earlier quoted context omitted.

I recently worked on information extraction from 10K documents. GPT-3 needs about 7 days of operation in batch mode on one thread. It takes 40..70s to read one single document and report the extracted data. One MINUTE per page. But I think you meant GPT-3 has seen many books during training, not during inference. You should know that training on millions of books is not the only way GPT-3 learns. It is just the found…

Your reply is mostly a red herring because as you yourself said, the OP was talking about regular training data, not the ICL you focused on. Just because you can find one way that GPT might be slow doesn't invalidate the point that its training does use massive amounts of data.

If GPT-3 can learn in context it means both the training set and the prompt could be in copyright violation. So even a clean model, trained on licensed data, cannot guarantee there will be no copyright issue.

Re: AI is in danger of being swallowed up by copyright law

#469
post #121

Earlier quoted context omitted.

AI companies can ask for permission if they want to train their models on other people's works Do you ask for permission when you train your mind on copyrighted books? Or observe paintings? Or listen to music? Do you ask for permission when you get new ideas from HN that aren't your own? Humans are constantly ingesting gobs of "copyrighted" insights that they eventually remix into their own creations without necessar…

> Do you ask for permission when you train your mind on copyrighted books? Or observe paintings? Or listen to music? Yes, that’s exactly what happens when you buy a book, or pay for a music subscription. The work is in the public domain, then global permission to observe and copy the work is already granted. > Do you ask for permission when you get new ideas from HN that aren't your own? You don’t need to. It’s impli…

> > Do you ask for permission when you train your mind on copyrighted books? Or observe paintings? Or listen to music?

> Yes, that’s exactly what happens when you buy a book, or pay for a music subscription. The work is in the public domain, then global permission to observe and copy the work is already granted.

You can buy a book, read it, sell the book, and then write and sell another book based on the ideas contained in the first book (Baker v Seldon). This is the cornerstone of contemporary copyright law. Or read the book on a shelf of a bookstore where the clerk is asleep. Or borrow the book from the library or any other manner where direct compensation of the author is nowhere to be seen.

Copyright is consistently interpreted in alignment with the needs of public learning, both by protecting the authorial incentive as well as protecting the public need for knowledge.

Re: AI is in danger of being swallowed up by copyright law

#470

Earlier quoted context omitted.

Doesn't follow. Why would a tool built in Japan limit itself to Japanese? ChatGPT doesn't limit itself to English.

To export you must abide to the laws of the country you're exporting to.

your novel theory of extraterritorial jurisdiction will no doubt be very interesting to trips litigators

afaik asahi v. superior court is still governing precedent in the usa though so it won't be of any interest to domestic litigators in the usa

Post reply on HN