Live data from Hacker News

Artificial Intelligence and Copyright: Request for comments

federalregister.gov

151–160 of 321 posts

Re: Artificial Intelligence and Copyright: Request for comments

#151

Much debate has been had about how existing copyright law applies to AI models. But once you get past that and start asking about how copyright should apply to AI models (as the copyright office is here) the answer in my mind becomes clear. Copyright, as defined in the U.S. Constitution, exists "to promote the Progress of Science and useful Arts"[1]. I can think of no better modern example of "the Progress of Science…

Do you see #1 and #3 conflicting at all? Ex: you produce a model, run it, publish and copyright some output. I can then use that as training data for another model in the style of your existing model?

Re: Artificial Intelligence and Copyright: Request for comments

#152
post #68

Earlier quoted context omitted.

I don’t think it makes sense for both model builders and the model’s users to separately obtain licenses for the same works used in the training set. A model trained on several copyrighted data sources cannot somehow be used in a way depending on a subset of those sources. So all parameters of usage and compensation should be settled by contract between the model builder and copyrighted data supplier, before the copy…

Yes, someone using a model can’t know if the generated text/image/sound is a nearly identical copy of the original material they don’t recognize. If use of the output of these systems comes at significant legal risk then then such systems become nearly useless.

> if the generated text/image/sound is a nearly identical copy of the original material they don’t recognize

how does the industry today deal with artists that "copy" off some other works? This isn't a problem with AI at all - just that AI provides a tool to generate such works faster.

Re: Artificial Intelligence and Copyright: Request for comments

#153
post #49

I believe we first need to answer the question of whether the copyright of the AI model’s source text or images affects the output. My opinion — and note I’m a software engineer, not a lawyer — is that an AI, being a statistical model and not generally intelligent, should not be allowed to disregard the copyright of its source material. This would, I think, require the AI’s creator to secure a license for all of its…

My problem with this is that artists learn by studying other artists, cutting that off because it's AI rather than focusing on whether the resulting work is derivative, seems more of a problem to me. It seems to me that an AI can be used for either original work or derivatives, proving that you can get derivatives out of it has always struck me as no different than commissioning a copy of someone's work from a human…

Can an AI express to you how van gogh affected it as an artist? I'm not sure that AI is "learning" the way we say humans are "learning," when humans learn and study art. Obviously there is no debate that you can input van gogh into a model and produce something van gogh-like as a result. But I've not seen anything that indicates that the AI is learning anything about van gogh at all. Perhaps it comes down to whether you think learning van gogh is just creating a mapping of all of his brush strokes ever, and only exactly what they look like. It's obvious the AI knows nothing more than that. If you think that's what humans do when they learn art, I'd be sad for you!

As to your hypothetical, we don't give copyrights to people who make rote copies of things, human or otherwise. Is the implication of the shock, that there is sufficient difference with the work as to render it a derivative and not a copy? Okay, how so? And of what consequence? Making derivatives of a copyright without license is infringement.

Re: Artificial Intelligence and Copyright: Request for comments

#154
post #148

Earlier quoted context omitted.

So i'm not sure how I feel, but to play Devil's advocate -- If I know anything I create is just going to be hoovered up and input into somebody's AI model so I do 99% of the work and they get 99% of the profit, perhaps I'm much less likely to progress Science and useful Arts by creating content in the first place. I fear an internet of signup walls and TOC agreements for everything, just to prevent crawlers that feed…

> If I know anything I create is just going to be hoovered up and input into somebody's AI model but today, without an AI model, anything you create is already going to be learnt and studied (if it is worth studying of course). What's the difference, but speed? > they get 99% of the profit Why is that a priori the assumption? What stops you from getting a profit? > I do 99% of the work you did 0.000001% of the work,…

> What stops you from getting a profit?

OpenAI and Stable Diffusion not paying for their dataset. I don’t believe GitHub asked for my contribution to Copilot.

Re: Artificial Intelligence and Copyright: Request for comments

#155
post #49

Earlier quoted context omitted.

My problem with this is that artists learn by studying other artists, cutting that off because it's AI rather than focusing on whether the resulting work is derivative, seems more of a problem to me. It seems to me that an AI can be used for either original work or derivatives, proving that you can get derivatives out of it has always struck me as no different than commissioning a copy of someone's work from a human…

You can ask someone to produce a pin-up version of Minnie Mouse, but good luck using it in any commercial activities. Most LLMs are just profiteering from people’s labor without their consent. And there’s nothing new being produced. It’s always a statistical output of previous works.

> a pin-up version of Minnie Mouse

that's not because of copyright, but because of trademark. If you make the minnie mouse sufficiently different that it cannot be mistaken for not being Minnie to the average person, and don't call it minnie mouse (to get rid of trademark), disney will have a much harder time suing you. Of course, they will still try, and steam roll you with just money instead.

Re: Artificial Intelligence and Copyright: Request for comments

#156
post #96

Earlier quoted context omitted.

Is intelligence really a factor here? Say I use the same training set as one of these LLMs, copyright protected text and all, and use it to derive a compression algorithm that uses very little space to store tokens and token sequences that are common in that huge collection of text. The resulting compression scheme includes some sort of statistical artifact derived from that copyrighted text. Is that allowed? And if…

Very good question indeed. A lot of these questions are somewhat ethical/moral in nature. E.g. is it okay to take someone else's creative work, process it through some algorithm, to create a service like ChatGPT? Or a compression algorithm? I don't know. It's awesome to see the Copyright office request input from both sides of the argument.

> is it okay to take someone else's creative work, process it through some algorithm, to create a service like ChatGPT? Or a compression algorithm?

and the test i use is: if they currently allow a human to perform this same task, then it is allowed to be done using an AI model.

Re: Artificial Intelligence and Copyright: Request for comments

#157
post #97

Earlier quoted context omitted.

Whether or not “humans do it” isn’t relevant. You can walk around with a copyrighted song in your head. That is not copyright infringement. But if you take that song, create a digital copy, and distribute it for money, then you are violating someone’s copyright. Additionally, our legal system requires a balance of probabilities. It’s hard to prove that someone was influenced by another work unless the similarities ar…

I challenge you to listen to 4 chords of awesome and tell me again about how every song is completely original. How does eragon exist when it's definitely ripped parts from star wars, etc...ai usually doesn't spit out a full plagiarism, but a loosely inspired work which is what most media we consume is. Edit: 4 chords of awesome link is https://youtube.com/watch?v=oOlDewpCfZQ&si=8vL6PbDnHiaffJh3

“Every song is completely original” is the opposite of what I said.

Re: Artificial Intelligence and Copyright: Request for comments

#158

I believe we first need to answer the question of whether the copyright of the AI model’s source text or images affects the output. My opinion — and note I’m a software engineer, not a lawyer — is that an AI, being a statistical model and not generally intelligent, should not be allowed to disregard the copyright of its source material. This would, I think, require the AI’s creator to secure a license for all of its…

I personally have a really hard time finding any meaningful difference or distinction between "AI" and "lossy compression". Copyright and "lossy compression" are pretty easy to reason about. Model "building" is "compression". Model "use" is "decompression". Everything about these AI models seems to be about the "lossy" part, but "lossy" is just an adjective to the main show. It's very difficult to not conclude that c…

Information is not copyrighted, just the expression of said information.

So if you took a recipe book, extracted the recipe information, and listed out the recipe in a different format (such as a table), it's a new work. It does not violate the copyright of the recipe book you extracted the info from.

Re: Artificial Intelligence and Copyright: Request for comments

#159

I'm going to try to plead my case for images generated using sophisticated prompt engineering to be copyrightable. For example, at the point that I've written a prompt with 20 tags, 10 negative prompt tags, some loras, custom weights, embeddings merges, and prompt editing, I'm now writing what is effectively a "program", which should be copyrightable and so should its outputs. It's total BS to me that a book of midjo…

The training model for Stable Diffusion has a lot of copyrighted images mixed together into an output which makes the plagiarism non-obvious, but let's reduce the set by 1 image. Shouldn't affect the output too much, right? Maybe some prompt will have a slightly different image. Now let's reduce it by another image. Again, less options for what to display, fewer images to take pixels from, but still a lot of options,…

It is impossible to 1:1 replicate the input as an output because the images are not stored. It isn't a database. It's basically aggregating summaries/abstractions/generalizations of a bunch of tags.

In other words, it is transformative by default.

Re: Artificial Intelligence and Copyright: Request for comments

#160
post #32

There are three copyright issues here; datasets, model weights, and model outputs. Dataset copyright is pretty well defined and things can often be used under fair use. Fair use decisions are done with a four prong test and really decided by the courts on a case-by-case basis. Model weights cannot currently be copyrighted. They are the output of a mechanical process over the dataset. However, software faced a similar…

Realistically an AI model is basically just a very complicated piece of software. The model weights are akin to the software code, the model outputs are akin to the outputs a user of the software creates, and the datasets are akin to the intellectual property put into the software by the developer to create the code.

In the same way that a developer could not simply steal someone elses intellectual property in order to develop a feature of a piece of software, one cannot simply steal the intellectual property to adjust the model weights. The main difference is its generally quite easy to see in practice if a model has utilized some intellectual property (because for example you can ask ChatGPT to recite the first 100 words of Harry Potter) compared to another piece of software where you'd need access to the source code or developers thoughts (which could only be achieved through litigation, in most circumstances).

I think a great many people come up with convoluted answers to this question because they are uncomfortable with the reality that these very large organizations have essentially stolen hoards of intellectual property, and now that the horse has bolted people want to justify not closing the barn door. It seems to me very simple: to train an AI model on data, you must respect its copyright. The model weights should be copywriteable by the developers of the model (even if the law currently does not allow this), and the outputs of the model should be copywriteable by the person who interacted with the model (software) to produce the outputs.

The analogy with Photoshop is extremely simple: If some other software invented Gaussian blurring and copywrited it, then Adobe would have to license that technology from them to include it as a feature in Photoshop. The actual photoshop software/code would be copywrited by Adobe, and if someone created an blurry image with Photoshop they can copywrite it.

I think people only disagree with this due to some sense that the process of translating data to model weights is "automatic" or "computational" in nature. You could in principle get a person to, by hand, go through millions of data sets and compute the changes to the model weights. This is no different to someone writing a piece of code, checking someone elses approach, and adjusting their own code after the fact. It just happens that we have developed very effective tooling to automate the adjusting of the code.

Post reply on HN