Live data from Hacker News

Artificial Intelligence and Copyright: Request for comments

federalregister.gov

181–190 of 321 posts

Re: Artificial Intelligence and Copyright: Request for comments

#181
post #32

There are three copyright issues here; datasets, model weights, and model outputs. Dataset copyright is pretty well defined and things can often be used under fair use. Fair use decisions are done with a four prong test and really decided by the courts on a case-by-case basis. Model weights cannot currently be copyrighted. They are the output of a mechanical process over the dataset. However, software faced a similar…

Realistically an AI model is basically just a very complicated piece of software. The model weights are akin to the software code, the model outputs are akin to the outputs a user of the software creates, and the datasets are akin to the intellectual property put into the software by the developer to create the code. In the same way that a developer could not simply steal someone elses intellectual property in order…

Three points I want to make:

- Models are nothing more than a statistical distillation of facts that can be traversed. They are not like software at all. Calling them software is like calling pachinko machines software. Nonsense.

- Models are mechanically derived with no element of human authorship or creativity. You could argue that there is creativity in selecting the dataset or the process that derives the model, but neither is relevant to the final generated model. Even if we assumed for the sake of argument that a model is more that a just statistical distillation, it should still not be considered copyrightable due to this reason alone.

- Don't use the word Steal when you refer to the well-defined act of infringement. Stealing implies deprivation of property which does not and cannot occur in this case. Using the word Infringe is more honest and less manipulative.

Re: Artificial Intelligence and Copyright: Request for comments

#182
post #106

I have never understood the fair use argument when it comes to training data. I publish a copyrighted article. Some LLM ingests it without permission, but since the output of that LLM is sufficiently different from my source article there is no violation. I publish copyrighted code. Some company decides to consume it without purchasing a license. The product they distribute is vastly different from my code itself, bu…

The copyrighted article is not part of the output (at least, not verbatim). The copyrighted code is part of the output.

Re: Artificial Intelligence and Copyright: Request for comments

#183

I believe we first need to answer the question of whether the copyright of the AI model’s source text or images affects the output. My opinion — and note I’m a software engineer, not a lawyer — is that an AI, being a statistical model and not generally intelligent, should not be allowed to disregard the copyright of its source material. This would, I think, require the AI’s creator to secure a license for all of its…

It shows you are not a lawyer. You misunderstand how copyright works. Creating copies or derivative works and distributing those is all that matters under copyright. This is not "disregarding" copyright (which is not an actual thing) but something that is either fair use or may require some kind of permission from the creators of the original by those distributing some kind of derived work or copy. That's why it's called copyright.

Copyright merely restricts the distribution of original works or their derivatives. In case of an infringement, copyright holders can insist you stop distribution and/or compensate them for that.

If I sell you a paint brush, I'm not liable for you putting a red nose on the mona lisa and trying to sell it off as an original work. Doing that on the original would be an act of vandalism (because you don't own it) and doing that on a replica that you got from somewhere infringes on the rights of those that created the replica. Which is a derived work or copy in itself of course and the distribution of that is regulated by copyright. Distribution of such a replica is of course fine because Da Vinci has been dead for a very long time and his work would no longer be protected under copyright. Distributing your red nosed mona lisa would therefore be fine too. Either way, the paint brush seller is no party in this case this is between you, Da Vinci, his descendants, and the replica creators.

Now your assertions as to what AIs are of aren't, are simply not relevant. You assert it's a statistics algorithms thingy. That sounds like a tool to me. Yet another paint brush. Using a paint brush is not infringing on anyone's rights. For that you have to distribute the results of your work. The nature of the tool does not matter. How you use the tool does not matter either. You merely create (potentially) derivative works with the tool and what you do with those matters. Especially when you distribute them to others. One of those derivative works is of course the AI model itself. Creating one is fine. Copyright gets potentially infringed when you distribute one.

Now we get to the core of the matter. Can you with a straight face say the AI model resembles the original and is a derivative work. It doesn't actually look like or resemble the original in any shape or form. Even proving the AI model is derived from the original is tricky. Copyright is not about protecting vague ideas or notions but the concrete shape or form of things. And it's only an infringement if you distribute a derived work or a copy of a thing to others. So, merely creating an AI model is not distributing anything to anyone. You are merely using tools to create something for yourself. An AI model in this case.

Distributing a verbatim copy of a book is an infringement. Citing the book in your own work is fair use (up to a point). Paraphrasing elements from the book, acknowledging it exists, taking inspiration of it, or reading it aren't copyright infringements.

The legal problem with AI models is that their concrete shape or form doesn't resemble the original inputs in any shape or form. Besides, companies like OpenAI don't actually distribute their AI models. They are huge; it's not very practical. They merely exploit those models to generate outputs to inputs from their users and customers. Are those outputs derivative works? Maybe, but that's where it gets tricky. They clearly aren't in the classical sense. Not even close. But if you somehow could conclude that they are, who is distributing that derivative work? Secondly, it the AI model is a tool, who actually creates those outputs and are those outputs protected under copyright? Who actually holds those rights? And how would you tell apart such an output from a human created one?

It's questions like this that make all this extremely murky from a legal point of view. IMHO without dramatic changes to copyright law or the way it has been commonly interpreted legally, it's just very poorly suited to do anything about stopping AI companies from doing what they are doing. You'd have to bend the conventional interpretation quite a bit for that. No doubt, there will be court cases where people will try to do that. But it will take many years before the dust settles on that. And I wouldn't get my hopes up on some unexpected/dramatic outcome.

Re: Artificial Intelligence and Copyright: Request for comments

#184
I wonder how they verify the personhood of the people making the comments. I can see this process being easily abused if the comments aren't taken by real people in person.

That said, I hope the US doesn't end up piling even more restrictions onto copyright. They'd only be shooting themselves in the foot. Copyright has completely failed to achieve the purpose for which it was intended to solve (only intended to give authors a short amount of time to profit off their efforts? look at it now). Perhaps it's time to rethink the concept of copyright as a whole before other countries beat the US to it.

Re: Artificial Intelligence and Copyright: Request for comments

#185
post #152
post #68

Earlier quoted context omitted.

Yes, someone using a model can’t know if the generated text/image/sound is a nearly identical copy of the original material they don’t recognize. If use of the output of these systems comes at significant legal risk then then such systems become nearly useless.

> if the generated text/image/sound is a nearly identical copy of the original material they don’t recognize how does the industry today deal with artists that "copy" off some other works? This isn't a problem with AI at all - just that AI provides a tool to generate such works faster.

The difference is the artists assertion that it’s either original or a copy from something else. DALLE 2 can’t tell you if it’s original or not. These AI’s have no idea and the company or group that created them doesn’t review individual output so they can’t say either.

Re: Artificial Intelligence and Copyright: Request for comments

#186
post #151

Much debate has been had about how existing copyright law applies to AI models. But once you get past that and start asking about how copyright should apply to AI models (as the copyright office is here) the answer in my mind becomes clear. Copyright, as defined in the U.S. Constitution, exists "to promote the Progress of Science and useful Arts"[1]. I can think of no better modern example of "the Progress of Science…

Do you see #1 and #3 conflicting at all? Ex: you produce a model, run it, publish and copyright some output. I can then use that as training data for another model in the style of your existing model?

I see that more as a conflict between #1 and #2, but fair point. In extreme cases, you could probably make a crude copy of a model by training a new model solely on the outputs of the first one. Normally that would be a derivative work, but that's inconsistent with the idea that training on copyrighted works is always permissible.

Maybe one way to resolve this would be to say there ought to be some practical limits on what percentage of the training data can be from any one individual source. If I train an model solely on the text of one book, for example, such that it's so overfitted that it can do nothing but regurgitate passages from that book, it's probably fair to call that a derivative work. The same would apply to a model trained solely on output from another model. (Though if it merely incorporates a few examples from a bunch of different models that would be okay.)

Re: Artificial Intelligence and Copyright: Request for comments

#187

Much debate has been had about how existing copyright law applies to AI models. But once you get past that and start asking about how copyright should apply to AI models (as the copyright office is here) the answer in my mind becomes clear. Copyright, as defined in the U.S. Constitution, exists "to promote the Progress of Science and useful Arts"[1]. I can think of no better modern example of "the Progress of Science…

> provided the output does not conflict with any preexisting copyright

I think this clause is doing all the work here and unfortunately in many cases there's no quick, automatic way to verify this.

Because of the legal costs involved in litigating who copied whom, I fear this would allow someone to sue the original creator of a work for infringing on the copyrighted output of an AI model trained on that work. If this seems far fetched, consider that this already happens with the DMCA and Creative Commons works: https://www.techdirt.com/2016/04/26/ifpi-files-dmca-takedown...

Re: Artificial Intelligence and Copyright: Request for comments

#188
post #82
post #56

Earlier quoted context omitted.

My opinion as a SWE who is dating a lawyer (joke, not a serious qualification but it does provide some insight): Generative models traverse and interpolate high dimensional state spaces. These state spaces are created from input data. I would argue people do the exact same thing - the first main difference is we can use novel inputs (e.g. we can use images or words to develop our music/temporal state spaces and vice…

I can appreciate how this line of thinking might be attractive. But IMO the human machine comparison doesn't lend itself much credence. We shouldn't assume that just because a human is allowed to do something, a machine is automatically allowed to do the same thing, too. I think some care should be taken when considering if we allow machines to have the same privileges as humans.

A machine is just a tool. It is the creator and the user of the machine that has the privileges he uses the tool with. I think we should be careful not to anthropomorphize, attribute agency, responsibility and autonomy to something that is essential a better photoshop plugin.

Re: Artificial Intelligence and Copyright: Request for comments

#189
post #164

Earlier quoted context omitted.

Someones comes to me to ask for a drawing of Batman or to write an erotic story around Supergirl. I can do it, but I cannot claim ownership over the characters. And I think I will quickly get a letter from DC or Marvel if I try to do this at scale.

> I can do it, but I cannot claim ownership over the characters. of course not. But you can claim ownership if you don't call those characters their original names, and make sufficient changes to the design (how sufficient is determined by a court of law - thus expenses). > DC or Marvel if I try to do this at scale. The show 'invincible'[1] has a character that is a basic copy of superman. And yet, you will find that…

> make sufficient changes to the design

I think that’s one of the issue. The transformation done by these tools are mechanical. Even if it may be extensive. The human input is too small. Omniman may have similarities with Superman, but he is not him in the larger context of the story. LLMs can not yet be that consistent for marketable output that deserves to be copyrightable.

I’m perfectly fine for LLMs to aid with spell checking and alternative phrasing (image is a grayer area). Bu the ideas of prompts and prompt output being copyrightable is something I oppose.

Re: Artificial Intelligence and Copyright: Request for comments

#190

Much debate has been had about how existing copyright law applies to AI models. But once you get past that and start asking about how copyright should apply to AI models (as the copyright office is here) the answer in my mind becomes clear. Copyright, as defined in the U.S. Constitution, exists "to promote the Progress of Science and useful Arts"[1]. I can think of no better modern example of "the Progress of Science…

So i'm not sure how I feel, but to play Devil's advocate -- If I know anything I create is just going to be hoovered up and input into somebody's AI model so I do 99% of the work and they get 99% of the profit, perhaps I'm much less likely to progress Science and useful Arts by creating content in the first place. I fear an internet of signup walls and TOC agreements for everything, just to prevent crawlers that feed…

> If I know anything I create is just going to be hoovered up and input into somebody's AI model so I do 99% of the work and they get 99% of the profit, perhaps I'm much less likely to progress Science and useful Arts by creating content in the first place.

Perhaps you wouldn't but I, and apparently most of the scientific community who publish research, would (and do). The entitlement of people here feel towards their way of doing things is astounding.

It's a story as old as time: Those who try to resist or limit progress (by placing arbitrary restrictions) will be beaten by those who adapt.

Post reply on HN