Live data from Hacker News

Artificial Intelligence and Copyright: Request for comments

federalregister.gov

261–270 of 321 posts

Re: Artificial Intelligence and Copyright: Request for comments

#261

Earlier quoted context omitted.

The thing you're authoring is just your prompt. Apply for a copyright for it. The image or text generated in reply is generated by a computer with no more input from you than someone commissioning work using very precise words, and lacks human authorship.

That's the same as saying if I open paint and draw with the mouse my mouse movements should be copyrightable but not the resulting image...

your mouse doesnt make decisions for you. ML based art does, which is why it lacks human authorship and you shouldn't be able to copyright it.

If you hand painted something in photoshop 100% you can copyright it. It has human authorship. If its mostly AI based fill, those elements can't be copyrighted. if Its 100% an ai result, its public domain.

Re: Artificial Intelligence and Copyright: Request for comments

#262

Earlier quoted context omitted.

I slightly disagree, in that I think the person using the tool should bear the burden of copyright. I.e. if the model outputs something under copywrite it merely can't be republished. In this same way, i can use Photoshop on proprietary data but I can't necessarily sell the results.

I see 2 problems with that. (1) how do you know if the image that just generated is substantially similar to an existing copyright work? Maybe if some registration tool existed, but other wise the burden is too great (2) what is stopping someone from generating millions of images and copy righting all the "unique" ones? Such that no one can create anything without accidental collisions.

> how do you know if the image that just generated is substantially similar to an existing copyright work?

This is already a problem with biological neural nets (i.e. humans). I remember as a teenager writing a simple song on the piano, and playing it for my mom; she said, "You didn't write that -- that's Gilligan's Island!" And indeed it was. If I had made a record and sold it, whoever owned the rights to the Gilligan's Island theme song could have sued me for it, and they would (rightly) have won.

There's already loads of case law about this; the same thing would apply to AI.

> what is stopping someone from generating millions of images and copy righting all the "unique" ones? Such that no one can create anything without accidental collisions.

Right now what's stopping it is that only humans can make copyrightable material; whatever is spat out from a computer is effectively public domain, not copyrighted.

Re: Artificial Intelligence and Copyright: Request for comments

#263
post #49

Earlier quoted context omitted.

My problem with this is that artists learn by studying other artists, cutting that off because it's AI rather than focusing on whether the resulting work is derivative, seems more of a problem to me. It seems to me that an AI can be used for either original work or derivatives, proving that you can get derivatives out of it has always struck me as no different than commissioning a copy of someone's work from a human…

You can ask someone to produce a pin-up version of Minnie Mouse, but good luck using it in any commercial activities. Most LLMs are just profiteering from people’s labor without their consent. And there’s nothing new being produced. It’s always a statistical output of previous works.

> You can ask someone to produce a pin-up version of Minnie Mouse, but good luck using it in any commercial activities.

The same would automatically apply to LLM output -- there's no need to change the current laws to cover that case.

The question is this. Suppose I ask a human artist and an LLM to create me a new female mouse cartoon character. And suppose both the artist and the LLM have been exposed to Minnie Mouse. It's not unlikely that the new character created in both cases will have aspects specifically similar to, or specifically opposite to Minnie Mouse.

In the case of the human artist, the new character will not be covered by Disney's copyright, unless there was a lot of copying. Why should the result be different for LLMs?

The logical conclusion of "any output of an LLM that's seen Minnie Mouse must be subject to Disney's copyright" is "any output of any human that's seen Minnie Mouse must be owned by Disney". Which I'm sure Disney would love, but would certainly make the world a worse place for everyone.

Re: Artificial Intelligence and Copyright: Request for comments

#264

I believe we first need to answer the question of whether the copyright of the AI model’s source text or images affects the output. My opinion — and note I’m a software engineer, not a lawyer — is that an AI, being a statistical model and not generally intelligent, should not be allowed to disregard the copyright of its source material. This would, I think, require the AI’s creator to secure a license for all of its…

I personally have a really hard time finding any meaningful difference or distinction between "AI" and "lossy compression". Copyright and "lossy compression" are pretty easy to reason about. Model "building" is "compression". Model "use" is "decompression". Everything about these AI models seems to be about the "lossy" part, but "lossy" is just an adjective to the main show. It's very difficult to not conclude that c…

> I personally have a really hard time finding any meaningful difference or distinction between "AI" and "lossy compression".

If you feed a photo of your dog into a JPEG compressor and the result looked like a cat in the same style, I think you'd be pretty annoyed.

Re: Artificial Intelligence and Copyright: Request for comments

#265

Earlier quoted context omitted.

The training model for Stable Diffusion has a lot of copyrighted images mixed together into an output which makes the plagiarism non-obvious, but let's reduce the set by 1 image. Shouldn't affect the output too much, right? Maybe some prompt will have a slightly different image. Now let's reduce it by another image. Again, less options for what to display, fewer images to take pixels from, but still a lot of options,…

This viewpoint is about as coherent as "every image file is copyright infringing because every pixel in it exists somewhere in some other image somewhere". Derivative works, when substantially changed, are not infringing. If I take an image of the Mona Lisa and rearrange all its pixels so it looks like a picture of a cat, that's not infringement. If I sample lines and curves and colors and styles from several images…

> If I take

> If I sample

You're not a computer program and your viewpoint is about as valid as "cars don't need speed limits because most humans can't run faster than 10mph and that speed is safe".

Copyright laws where made with humans in mind.

Re: Artificial Intelligence and Copyright: Request for comments

#266
post #145

Earlier quoted context omitted.

Some compression, yes, but the analogy oversimplifies. AI rerepresents input information in a transformative way (embedding, say) then creates new, derived and combined output from a new input (e.g prompt). It's not just lossy compression. It's potentially novel.

It’s compression + filtering. Nothing generative. Its output is like 99.99 % deterministic.

Source?

Re: Artificial Intelligence and Copyright: Request for comments

#267
post #140

Earlier quoted context omitted.

A divergence, but I see a lot of posters asserting that "humans learn by copying other people, but we don't call that a violation of copyright when they draw" People casually asserting that software is equivalent to humanity will be a non-negligible thing to consider, as irritating and poorly-founded as it seems. If the reproduction isn't pixel-perfect, but merely obvious and overwhelming, how do you refute that phil…

A human is still entering the prompt to generate the possibly copyrighted image/text. I don't think copyright law should care about the implementation. It's ok to copy a style if you use paint brushes or photo shop. But not ok if you use a statistic model?

Apply for a copyright on your human authored prompt then. That's the extent of human authorship.

Re: Artificial Intelligence and Copyright: Request for comments

#268

Earlier quoted context omitted.

The training model for Stable Diffusion has a lot of copyrighted images mixed together into an output which makes the plagiarism non-obvious, but let's reduce the set by 1 image. Shouldn't affect the output too much, right? Maybe some prompt will have a slightly different image. Now let's reduce it by another image. Again, less options for what to display, fewer images to take pixels from, but still a lot of options,…

> Once there are millions of images, you end up with just tiny blobs of pixels being copied from many different images. This is not how these neural nets work. They don't copy pixels from anywhere. They learn features. The features represented internally are generally not easy to interpret to humans, but for sake of illustration, there could be an artificial neuron that fires when a subject should have blue eyes. Hav…

Thank you for the explanation. Let me explain my position in similar terms.

I'm not replicating an image, I'm "using my brain to build a network of neurons that map electrical impulses from the optical nerve excited by wavelengths projected onto my retina in order to send other electrical signals to actuator tissues".

The complexity of the process is irrelevant imo. We can treat it as a black box and look at the inputs and outputs.

If the images in the database didn't exist, it wouldn't know what to draw, and those images are copyrighted.

Everyone's welcome to take a camera, run around the world and label every object for the neural net to learn, like a human does, but model authors didn't do that because using copyrighted images for free is much easier.

Re: Artificial Intelligence and Copyright: Request for comments

#269

Earlier quoted context omitted.

The training model for Stable Diffusion has a lot of copyrighted images mixed together into an output which makes the plagiarism non-obvious, but let's reduce the set by 1 image. Shouldn't affect the output too much, right? Maybe some prompt will have a slightly different image. Now let's reduce it by another image. Again, less options for what to display, fewer images to take pixels from, but still a lot of options,…

It is impossible to 1:1 replicate the input as an output because the images are not stored. It isn't a database. It's basically aggregating summaries/abstractions/generalizations of a bunch of tags. In other words, it is transformative by default.

Can it replicate the input 0.1:0.1?

Re: Artificial Intelligence and Copyright: Request for comments

#270
post #255
post #210

Earlier quoted context omitted.

> DALLE 2 can’t tell you if it’s original or not whoever pressed the button to run DALLE will make the assertion, just like whoever was running photoshop to make the image today would make the same assertion.

Based on what? A photoshop user controls what data photoshop uses, a DALLE user doesn’t. Even a prompt as generic as “Cat” could be producing an obviously derivative work if you compare it to the original. This is true for all prompts.

> A photoshop user controls what data photoshop uses

the point was that the user of the program is making their declaration, whether it's photoshop or DALLE. How does the business verify that their staff artists aren't producing copyright infringing material, just from memory?

The liability falls to them to verify the copyright status of the output they're asked to make. A business paying a photoshop user to produce a picture has just as much (or as little) trust in them as the button presser for DALLE.

Post reply on HN