Live data from Hacker News

AI is in danger of being swallowed up by copyright law

heathermeeker.com

241–250 of 705 posts

Re: AI is in danger of being swallowed up by copyright law

#241

There's no part of AI that is being swallowed up by copyright. AI companies can ask for permission if they want to train their models on other people's works. It's not that hard, various image hosting sites have already added an opt-in/opt-out toggle to their services. Sites might even get away with using this stuff as compensation for free hosting. The fact of the matter is that the AI companies don't want to ask fo…

> but what the customers of AI people want isn't available under those terms. What the customers of AI want is accurate predictions of the models, and they can get that even if everyone demanding to get removed from the training set would be removed. The makers of generative AI could remove every living artist who wants to from the dataset, the model would still develop a general solution of color theory, composition…

Then do it!

I swear when I see this argument because it makes me angry.

You’re right, but they didnt, because they were too lazy and cheap to do it that way.

…and that’s why people are angry, and rightly so. Fully licensed models are the future, and it’s both irritating and disappointing that we are where we are right now because the people training these models were too lazy to assemble a training dataset that wasn’t problematic (ie. full of porn and copyrighted material).

You can argue the “but at the end of the day it’s all the same…” argument if you like, but clearly from the lawsuits it isn’t ok

They’ve completely messed it up.

There’s a reason the openai api terms of service says that “the Content may be used to improve and train models”; they’re setting themselves up to have a concrete defence for the source training data for their models.

Good job.

Stability can burn in a fire. They’ve really trashed the reputation of generative AI in a way that is going to be very difficult to recover from.

Re: AI is in danger of being swallowed up by copyright law

#242
post #153

Earlier quoted context omitted.

> “Do you ask for permission when you train your mind on copyrighted books?” I’m not able to read billions of books in less than an hour. Even if we agree that machine learning is like human learning, scale commonly matters in law.

I recently worked on information extraction from 10K documents. GPT-3 needs about 7 days of operation in batch mode on one thread. It takes 40..70s to read one single document and report the extracted data. One MINUTE per page. But I think you meant GPT-3 has seen many books during training, not during inference. You should know that training on millions of books is not the only way GPT-3 learns. It is just the found…

Your reply is mostly a red herring because as you yourself said, the OP was talking about regular training data, not the ICL you focused on.

Just because you can find one way that GPT might be slow doesn't invalidate the point that its training does use massive amounts of data.

Re: AI is in danger of being swallowed up by copyright law

#243

I'm hooked on Stable Diffusion. It's the most impressive sudden leaps in technology in my 30 or so years of being old enough to understand it. Much of the power it has, is based on the material it is trained on, and I'm very appreciative of the work and effort that was made in order to make it what it is. And, it's only getting started. That said... I get it. Huggingface are working on a diffusion based model for mus…

Music business has for reasons connected to broadcast, phonograph and CD been extremely profitable. Less in the streaming world but still ok. This powers a lot of lawyers for the next decades. Visual artists earned a lot less.

Re: AI is in danger of being swallowed up by copyright law

#244
post #121

There's no part of AI that is being swallowed up by copyright. AI companies can ask for permission if they want to train their models on other people's works. It's not that hard, various image hosting sites have already added an opt-in/opt-out toggle to their services. Sites might even get away with using this stuff as compensation for free hosting. The fact of the matter is that the AI companies don't want to ask fo…

AI companies can ask for permission if they want to train their models on other people's works Do you ask for permission when you train your mind on copyrighted books? Or observe paintings? Or listen to music? Do you ask for permission when you get new ideas from HN that aren't your own? Humans are constantly ingesting gobs of "copyrighted" insights that they eventually remix into their own creations without necessar…

> Do you ask for permission when you train your mind on copyrighted books?

Yes, you do need to buy books, which gives you permission to read them.

Re: AI is in danger of being swallowed up by copyright law

#245
post #228
post #177

Earlier quoted context omitted.

I’m not able to read billions of books in less than an hour. I think you underestimate the sheer volume of data + conclusions the brain ingests and processes on a daily basis, primarily through unconscious experience.

The training set for GPT-3 is about 500e9 tokens; any given synapse in a human in their lifetime is going to fire about 2e9s * (10% * {100Hz to 1000Hz}) = 20e9 to 200e9 times.

It sounds like you agree with them, since there are a lot of synapses.

Re: AI is in danger of being swallowed up by copyright law

#246
post #235

Earlier quoted context omitted.

[flagged]

I don't think it will benefit "the little guys", because "the little guys" rarely have the resources and time to litigate in the first place, or lobby to lawmakers to make the details work in their favour. Copyright always benefits "the big corps". Everyone is so eager to get one up on Microsoft that they're forgetting the bigger picture.

The fact that the justice system is so inefficient that it doesn't serve people with less than a lawyers salary of money to waste isn't important to the conversation of whether it's fair or not for someone to ignore licensing and repackage your code as an "AI".

If you want to talk about the big picture here, it's about privatizing gains and socializing losses, the goal of every bigcorp, which is just more reason to disallow this abuse.

Re: AI is in danger of being swallowed up by copyright law

#247
post #235

Earlier quoted context omitted.

[flagged]

I don't think it will benefit "the little guys", because "the little guys" rarely have the resources and time to litigate in the first place, or lobby to lawmakers to make the details work in their favour. Copyright always benefits "the big corps". Everyone is so eager to get one up on Microsoft that they're forgetting the bigger picture.

Thank you. It’s legitimately scary to read through these comments. It’s like watching everyone clamor for Stalin to be put in power: not even a good idea in the short term, let alone the long term.

Re: AI is in danger of being swallowed up by copyright law

#248
post #121

There's no part of AI that is being swallowed up by copyright. AI companies can ask for permission if they want to train their models on other people's works. It's not that hard, various image hosting sites have already added an opt-in/opt-out toggle to their services. Sites might even get away with using this stuff as compensation for free hosting. The fact of the matter is that the AI companies don't want to ask fo…

AI companies can ask for permission if they want to train their models on other people's works Do you ask for permission when you train your mind on copyrighted books? Or observe paintings? Or listen to music? Do you ask for permission when you get new ideas from HN that aren't your own? Humans are constantly ingesting gobs of "copyrighted" insights that they eventually remix into their own creations without necessar…

> Do you ask for permission when you train your mind on copyrighted books?

I pay the books directly (cash, credit) or indirectly (school books via taxes). I do pay the louvre to observe the painting. I also pay to listen music in ads (YouTube) or via subscription (YT Music and Spotify).

Re: AI is in danger of being swallowed up by copyright law

#249
post #116

Earlier quoted context omitted.

It’s strange to me that there’s a lot of overlap between people who think AI training should require explicit consent for every piece of training data, and people who think copyright and patents are insanely restrictive in the music/movies/literature/software world. It’s also worrying that requiring consent to train an AI model will inevitably lead to requiring consent to make handmade art that’s a little too similar…

All or nothing, in my opinion. Either abolish or severely reduce copyright, or abide by it. The simple fact of the matter is that Disney and Getty invest a lot of money into these materials being out there in the first place. Open source programmers and artists spend a lot of time producing works for no cost other than some minor courtesies. AI companies aren't your friend or the little mom ''n pop shop down the road…

> All or nothing, in my opinion. Either abolish or severely reduce copyright, or abide by it.

I firmly believe in "practice what you preach". I you declare you firmly believe in A but then do something directly counter to that because it's more convenient in this specific case, then that doesn't sit right with me.

Besides, further expanding copyright in this one area will only make it so much harder to reduce it later. And the pro-copyright folks will be able to say "you say you want less copyright, but you vigorously advocated in favour of copyright then, you hypocrite!" (and they wouldn't be entirely wrong, either). All this effort and energy fighting ML tools would be better directed at reducing copyright instead.

I don't disagree with your view on corporations. Do I like what CoPilot is doing? Not really. But at the end of the day: does CoPilot's or ChatGPT's mere existence really take away anything concrete from me? Am I harmed or even inconvenienced by it? Are my rights reduced? Is my code harmed by it? Is my income reduced? I don't really see how it concretely affects me, other than a general "feeling of unfairness".

And I see real risks with all of this: most regular people and small businesses don't have the resources to litigate as it's expensive and time-consuming, so a "license" that you or I slap on a piece of code is, realistically speaking, just ink on a piece of paper. GPL violations are rampant, violations of other licenses probably happen even more (but people generally care less about that, so not as widely publicized). Who will benefit with more copyright law on their side? The ones with deep pockets and many lawyers on retainer. i.e., the corporations neither of us like. Think creative new copyright lawsuits such "we claim copyright on the Java API" kind of stuff.

Re: AI is in danger of being swallowed up by copyright law

#250
post #232

I think the underlying question is one of "degrees of derivation". There's a famous Carl Sagan quote: “If you wish to make an apple pie from scratch, you must first invent the universe” which hints at the problem: Nothing is created in a vacuum. Let's compare what Stable Diffusion does with what Franz von Holzhausen, head of design at Tesla, does. Franz didn't come into existence out of nothing and knew how to design…

The key point is: an AI calls itself an intelligence but this so far is quite a bit of marketing. A human is considered an intelligent being. Where a process of creating new things happen. They are based on true learning and not just reproduction. And where they have been just reproduction, of course it went to the courts.

I'm really unsure if there is a qualitative difference between a human looking at lots of images, deriving patterns and recompiling them into a new image or a computer doing the same.
Post reply on HN