Live data from Hacker News

Artificial Intelligence and Copyright: Request for comments

federalregister.gov

21–30 of 321 posts

Re: Artificial Intelligence and Copyright: Request for comments

#21

I'm going to try to plead my case for images generated using sophisticated prompt engineering to be copyrightable. For example, at the point that I've written a prompt with 20 tags, 10 negative prompt tags, some loras, custom weights, embeddings merges, and prompt editing, I'm now writing what is effectively a "program", which should be copyrightable and so should its outputs. It's total BS to me that a book of midjo…

The training model for Stable Diffusion has a lot of copyrighted images mixed together into an output which makes the plagiarism non-obvious, but let's reduce the set by 1 image. Shouldn't affect the output too much, right? Maybe some prompt will have a slightly different image.

Now let's reduce it by another image. Again, less options for what to display, fewer images to take pixels from, but still a lot of options, output may sound copyrightable.

Now lets do that N-1 times. What output will we get when the model was trained on a single image, let's say an image that is labeled 'dog'. If your prompt is "an image of a dog" you will get that image, the only image in the training set. When going from latent space to image space, taking pixels from that image in the output, despite it being done in convoluted ways, is that not an obvious copyright infringement? I think it is. There's a cloud of mumbo jumbo about latent space, but after the dust settles and it needs to generate pixels in the output image, Stable Diffusion has a step that is essentially copying pixels from the source image into the output. When there's only 1 image, it will reproduce large portions of that image, necessarily infringing on copyright.

So then adding back images one by one into the training set, each one being used as source for the pixels being copied, what makes that model OK? Just because the output is 50% image A and 50% image B, or 0.1% image A and 0.1% image B and 99.8% image C, doesn't suddenly make it OK.

Once there are millions of images, you end up with just tiny blobs of pixels being copied from many different images. That's still infringes on the copyright of all those images, because it's essentially a map-reduce process that maps pixels from copyrighted images and reduces them into a single image.

Re: Artificial Intelligence and Copyright: Request for comments

#22

Earlier quoted context omitted.

Thats nice and all. But it has nothing to do with whether something is copyrightable or not. Instead, something is copyrightable based on the amount of human input into the process. And it is very clear that there can be a lot of human input into AI image generation. Even though I will concede that going into midjourney and just typing in "Hot anime girl" isn't a lot of input and likely doesn't deserve copyright prot…

No your human input product is the inputs you authored but you're applying for copyright for something else.

So then yes, it is about the human input into the process. Thats what I just said.

Having large amounts of human input is the thing that matters for this stuff, which is the case for many forms of AI art.

In the same way how photoshop uses a computer, and the computer creates the art, the resulting computer generate art can still have copyright protections. (because of the large amount of human input, even though yes it used a computer)

Re: Artificial Intelligence and Copyright: Request for comments

#23

I believe we first need to answer the question of whether the copyright of the AI model’s source text or images affects the output. My opinion — and note I’m a software engineer, not a lawyer — is that an AI, being a statistical model and not generally intelligent, should not be allowed to disregard the copyright of its source material. This would, I think, require the AI’s creator to secure a license for all of its…

Is intelligence really a factor here?

Say I use the same training set as one of these LLMs, copyright protected text and all, and use it to derive a compression algorithm that uses very little space to store tokens and token sequences that are common in that huge collection of text. The resulting compression scheme includes some sort of statistical artifact derived from that copyrighted text. Is that allowed? And if so why is an LLM different?

Re: Artificial Intelligence and Copyright: Request for comments

#24

The narrative beauty of filling this up with GPT-created comments will be so incredibly sublime.

Since it'll be real people driving the efforts; the final results should normalize out about the same, statistically speaking.

This all gets boiled down to like 3 or 4 rows in a spreadsheet, anyway.

Re: Artificial Intelligence and Copyright: Request for comments

#25
post #7

Earlier quoted context omitted.

No, you're passing inputs to a program, and your description completely omits the vast majority of that input: the creative output of an unknown number of other people , whose rights you are attempting to launder.

Thats nice and all. But it has nothing to do with whether something is copyrightable or not. Instead, something is copyrightable based on the amount of human input into the process. And it is very clear that there can be a lot of human input into AI image generation. Even though I will concede that going into midjourney and just typing in "Hot anime girl" isn't a lot of input and likely doesn't deserve copyright prot…

If the court is convinced that prompt engineering is "original and creative".[1]

Maybe it is. Or maybe it's more like tweaking the random seed.

[1]: https://www.copyright.gov/comp3/chap300/ch300-copyrightable-..., see "The Originality Requirement" and "Creativity".

Re: Artificial Intelligence and Copyright: Request for comments

#26

We as a society have a relatively healthy setup for people to create art and content. Sure there are problems, but on the whole it mostly works. What AI will do is destroy that by removing the profitability of creating that content. Although generative AI operates on a similar principle to a human being exposed to a large number of artworks, it does so at a speed blindingly faster, enabling it to outcompete humans at…

>And I sincerely hope they are used against AI to make AI unprofitable.

No, they'll make AI unprofitable for small time creators but not massive corporations. The latter either already have rights to vast quantitates of training data or will hire a thousands in Africa to create training data that is just legally different enough to count.

Re: Artificial Intelligence and Copyright: Request for comments

#27

I believe we first need to answer the question of whether the copyright of the AI model’s source text or images affects the output. My opinion — and note I’m a software engineer, not a lawyer — is that an AI, being a statistical model and not generally intelligent, should not be allowed to disregard the copyright of its source material. This would, I think, require the AI’s creator to secure a license for all of its…

> is that an AI, being a statistical model and not generally intelligent, should not be allowed to disregard the copyright of its source material

None of what you are saying has anything to do with copyright.

The tool Photoshop isn't generally intelligent either. And yet, yes it can be used to create art using other people's stuff.

And it could be done legally if the results are transformative.

Re: Artificial Intelligence and Copyright: Request for comments

#28

I would say, treat the AI like a human viewing said data/material, but unlike the majority of humans, most dont have Kim Peek levels of observation and recall, so should there be some sort of expiration of data built into AI's and if so, should it be a blanket cut off date for all data an AI has been exposed to, or allow for some specialisation like a human might have in order to fulfil their occupation? I'm also awa…

It's funny to see this copyright overreach as somehow opposed to monopolies. Should using data to train AI models be made illegal then the only entities capable of affording it would be global monopolies (and maybe pirates? Though the compute necessary may be prohibitive). There would be no way for smaller AI entities to compete.

Re: Artificial Intelligence and Copyright: Request for comments

#29

We as a society have a relatively healthy setup for people to create art and content. Sure there are problems, but on the whole it mostly works. What AI will do is destroy that by removing the profitability of creating that content. Although generative AI operates on a similar principle to a human being exposed to a large number of artworks, it does so at a speed blindingly faster, enabling it to outcompete humans at…

Even if we keep draconian copyright laws they need to be changed in some ways. Otherwise Disney will dominate by creating their own AIs and we still end up with a small faction controlling what we consume - and we will be paying them for the honor of it on top.

Re: Artificial Intelligence and Copyright: Request for comments

#30

We as a society have a relatively healthy setup for people to create art and content. Sure there are problems, but on the whole it mostly works. What AI will do is destroy that by removing the profitability of creating that content. Although generative AI operates on a similar principle to a human being exposed to a large number of artworks, it does so at a speed blindingly faster, enabling it to outcompete humans at…

>And I sincerely hope they are used against AI to make AI unprofitable. No, they'll make AI unprofitable for small time creators but not massive corporations. The latter either already have rights to vast quantitates of training data or will hire a thousands in Africa to create training data that is just legally different enough to count.

That is why we should halt AI completely and do a more thorough analysis of its societal-level implications before blindingly putting it out there.

Because when new technology is introduced, it makes it almost impossible to stop using it due to the way our current society is setup (as a sensitive machine that is very quick to reward any gains in efficieny and economic output as opposed to sustainability).

Post reply on HN