Earlier quoted context omitted.
Yeah, I think we are at the point where copyright doesn't exist anymore, at least for AI
All of human knowledge (an exaggeration, I know) at our finger tips. It's the most punk rock, anarchist thing tech has done since the internet and it's funny it's shaped as a product.
AI is just unauthorised plagiarism at a bigger scale
321–330 of 783 posts
Re: AI is just unauthorised plagiarism at a bigger scale
#322Earlier quoted context omitted.
This is a strictly worse world in almost every sense. It's as if we abolished physical property rights and suggested people arm themselves to keep what is (was) theirs instead. Civilization, gone.
It’s a false equivalence to say that intellectual property is property. Taking your car deprives you of your car. Taking your idea lets civilization advance.
Re: AI is just unauthorised plagiarism at a bigger scale
#323Earlier quoted context omitted.
This is naive in the opposite. Creators gonna create.
Who is giving a creator millions of dollars to create something if there is no guaranteed path to recouping production costs. Are we going the communist soviet union route where everything is decided by central committee?
Those of us who create for creation's sake need no other reason. I create because I want to, not because I want to use it to gain capital.
Sure, those lines get muddy when you want to do it professionally, but that's a separate argument.
Re: AI is just unauthorised plagiarism at a bigger scale
#324Re: AI is just unauthorised plagiarism at a bigger scale
#325Earlier quoted context omitted.
> I'd like to understand why I can't use a song in one of my videos without permission/payment, but an AI company can train models using that song without having either. Because training isn't redistribution. You can also listen to the song and make a new one that sounds similar, just like the AI can.
To do that training, you must first obtain the item with the content you require. Did OpenAI purchase a copy of every book they trained their models on? Answer: They did not. That is literally why there are dozens of ongoing lawsuits in progress.
Re: AI is just unauthorised plagiarism at a bigger scale
#326Earlier quoted context omitted.
I appreciate your comment, but you answered as if this question had been answered legally. It has not. The New York Times is suing both OpenAI and Microsoft for copyright infringement. The Authors Guild is suing OpenAI. Getty Images is suing Stability AI. Disney is suing Midjourney. Universal Music Group and Sony have filed suits against multiple AI companies. > so copyright doesn’t get involved at all. The dozens of…
Which statement of mine do you think is not settled law? Which law do you think is being broken and how? Your objection doesn’t make sense. In the event that an AI company loses a lawsuit for copyright infringement based on simply training on copyrighted works, the answer to you saying you’d like to understand why they can do it and you can’t is simply “your premise is wrong; neither of you can” .
I object to your statement that "copyright doesn’t get involved at all" when that is objectively untrue. If that was true, many of the world's largest companies wouldn't be spending tens of millions of dollars to have that question answered in court. Go to any law-focused forum, and you will find attorneys arguing over these questions.
To train a model using a book, you must first obtain a copy of that book. Did OpenAI purchase a copy of every book not already in the public domain used during training? They did not.
Some of the suits I mentioned claim that OpenAI literally stole copies of books to train its models.
My point is that the copyright question has not been answered. If the NYT, et. al. win, it will be a watershed moment for how AI companies pay for training data moving forward.
Re: AI is just unauthorised plagiarism at a bigger scale
#327Earlier quoted context omitted.
Can you explain how something like the Lord of the Rings film series gets created in a world with no IP laws.
Many versions are made, the best ones get the most views. You don't need huge budgets and guaranteed revenue to make great art. In fact, I'd argue it's often the opposite. Most big budget movies suck these days.
Re: AI is just unauthorised plagiarism at a bigger scale
#328IP attorney here and actively working on this problem. nla: if you create content online (public repo code, blog, podcast, YouTube, publishing) the smartest thing you can do if to file a US copyright, even if you have a hobby blog. Anthropic paid $1.5B in a class settlement to authors because it was piracy of copyrighted works. If we as a HN community had our works protected, there are potentially huge statutory dama…
Wait what do you mean by "file a copyright"? I have never heard of this, all explanations of copyright I have heard say that you automatically own the copyright to the things you make; and that "all rights are reserved" by default unless you give up on them through granting a license. Is this no longer the case? Why is this now suddenly different? When did it change?
There are tens of millions of registered copyrights in the US, nearly every published book, music, artwork, many magazines and major websites. Here's the official link, you can search the registry and there is a ton of info: https://www.copyright.gov/registration/
Re: AI is just unauthorised plagiarism at a bigger scale
#329Earlier quoted context omitted.
This is an incredibly naive view of intellectual property. If you cannot own things you create, there is little incentive to create and share those things. Do you think any of your favorite movies and TV shows ever get made without copyright protections? Of course not, because money needs to change hands for those things to be funded.
Yes, absolutely, and that is why history shows so few examples of any art having been created prior to the invention of copyright: nobody had any reason to do it.
Re: AI is just unauthorised plagiarism at a bigger scale
#330This is really not so clear cut as "fair use" might cover 99% of all data scrapping; you are not reproducing the originals just use them to estimate probabilistic distribution of tokens in pre-training. You are never going to get the exact book word-for-word using LLMs.
I don’t buy this argument. The tokens are useless without their context, which provides the probability distributions needed to make them useful. Sure you MIGHT not be able to get the book word for word, but it’s impossible to make a useful model without the whole book and all of the artistry that went into it, to guide the tokens in their expected output. Fair use generally does not cover commercial use, which this…
Commercial use counts _against_ a fair use defense, but is not dispositive: it's not accurate at all to say it "generally does not cover" commercial use. This is the "purpose and character" test, one of four in contemporary (United States) fair use doctrine.
Purpose and character also includes the degree to which a use is _transformative_. It's clear that the degree to which a training run mulching texts "transforms" them is very high. This counts toward a fair use finding for purpose and character.
> is dependent on the amount of the original content present in the derived work, which I would contend in this case is “all of it”
The "amount and substantiality" test. Your case for "all of it" can't possibly be sustained: the models aren't big enough. It's amount _and_ substantiality: this has come up in the publication of concordances, where a relatively large amount of a copyrighted work appears, but it's chopped up and ordered in a way which is no longer substantially the same. Courts have ruled that this kind of text is fair use, pretty consistently. It's not an LLM, of course, but those have yet to be ruled on.
Also worth knowing that courts have never accepted reading or studying a work as incorporation, and are unlikely to change course on the question. It's taken for granted that anyone is allowed to read a copyrighted work in as much detail as they wish, in the course of producing another one. Model training isn't reading either, but the question is to what degree it resembles study. I'd say, more than not.
Specifically:
> it’s impossible to make a useful model without the whole book and all of the artistry that went into it
Courts have never once accepted "it would be impossible for defendant to write his biography without reading plaintiff's" as valid, and it's been tried. The standard for plagiarism is higher than that.
"Effect upon the work's value" is probably the most interesting one. For some things, extreme, for others, negligible. I suspect this is the one courts are going to spend the most time on as all of these questions are litigated.
Ultimately, model training is highly out-of-distribution for the common law questions involving fair use. It was not anticipated by statute, to put it mildly. The best solution to that kind of dilemma is more statute, and we'll probably see that, but, I don't think you'll be happy with the result, given what I'm replying to. Just a guess on my part.