Earlier quoted context omitted.
> I've seen a lot of confidence on HN and other tech communities that a court would never rule that training an AI on copyrighted images is infringement, but I'm not so sure. To be clear, I hope that training AI on copyrighted images remains legal, because it would cripple the field of AI text and image generation if it wasn't! Regardless of the copyright of the training data which really is unresolved, the copyright…
AI-produced art is still human-made, as a person does the job of engineering a prompt and selecting from the generated images. The copyrightability of such work is unlikely to ever seriously be in question.
Getty Images bans AI-generated content over fears of copyright claims
211–220 of 390 posts
Re: Getty Images bans AI-generated content over fears of copyright claims
#212Earlier quoted context omitted.
Stable Diffusion is already out in the world. The cat is out of the bag.
Yah, but think of Napster getting eventually usurped by Spotify. The danger is that it's legally no longer possible to update the models (which are very expensive to train), and we end up with only Disney with the copyright horde large enough to train decent models, let alone good ones...
Re: Getty Images bans AI-generated content over fears of copyright claims
#213Earlier quoted context omitted.
> I've seen a lot of confidence on HN and other tech communities that a court would never rule that training an AI on copyrighted images is infringement, but I'm not so sure. To be clear, I hope that training AI on copyrighted images remains legal, because it would cripple the field of AI text and image generation if it wasn't! Regardless of the copyright of the training data which really is unresolved, the copyright…
AI-produced art is still human-made, as a person does the job of engineering a prompt and selecting from the generated images. The copyrightability of such work is unlikely to ever seriously be in question.
Re: Getty Images bans AI-generated content over fears of copyright claims
#214It's always weird to see the contrast between HN's reaction to copyright questions about text/image generation, and HN's reaction when it's code generation. When a model is trained on 'all-rights-reserved' content like most image datasets, the community say it's fair game. But when it's 'just-a-few-rights-reserved' content like GPL code, apparently the community says that crosses a line? Realistically, this tells me…
It's possible for generation models to perfectly memorize and reproduce training data, at which point I view it as a sort of indexed slightly-lossy compression, but it's almost never happening with image generation because the models are too small to memorize billions of pictures, it can't produce copies.
Stable diffusion 1.4 has been shrunk to around 4.3GB, and has around 900 million parameters.
I don't know how big Copilot is, but a relatively recently released 20 billion parameter language model is over 40GB. ( https://huggingface.co/EleutherAI/gpt-neox-20b/tree/main ) GPT-3, according to OpenAI, is 175 billion parameters.
It's possible there are some images in there you can pull out exactly as is from the training data, if they were to appear enough times, like I suspect the Mona Lisa could be almost identically reconstructed, but it would take a lot of random generation. I'm trying it now and most of the images are cropped, colors blown out, wrong number of hands or fingers, eyes are wrong, etc.
Re: Getty Images bans AI-generated content over fears of copyright claims
#215Earlier quoted context omitted.
Why would you search a Getty competitor for AI generated images when you can just roll your own?
It's faster to sift through pre-generated images than to build novel ones.
Re: Getty Images bans AI-generated content over fears of copyright claims
#216It's always weird to see the contrast between HN's reaction to copyright questions about text/image generation, and HN's reaction when it's code generation. When a model is trained on 'all-rights-reserved' content like most image datasets, the community say it's fair game. But when it's 'just-a-few-rights-reserved' content like GPL code, apparently the community says that crosses a line? Realistically, this tells me…
If this was a website for artists instead of programmers you'd see the exact opposite pattern. Unfortunately, people only seem to care when it threatens their own livelihood, not when it threatens that of the people around them.
An AI can probably do half of my day job because it's stupidly repetitive. Leadership imposes old ways, and they dismiss anything "new" (i.e. newer than 2005). For example, writing all these high-level data pipelines and even web backends in C++ despite having no special performance need for it. Even though I'm not literally copy-pasting code, I'm copy-pasting something in my mind only a little higher-level than that, then relying on some procedural and muscle memory to pump it out. It's a skill that anyone can learn, just takes time. If I didn't have side projects, I'd forget what it's like to think about my code.
Some old-school programmers complain about kids with high-level languages doing their job more efficiently, so they work on lower-level stuff instead. It's been that way for decades. Now AI is knocking on that door. But before AI, C was the high-level thing, and we got compiler optimizations obviating much of the need for asm expertise, undoubtedly pissing off some who really invested in that skillset. If I'm working on something needing the performance guarantees of C or Asm, and the computer can assist me, I'm all for it. Please take this repetitive job so I can use my brain instead.
And the copyright thing is just an excuse. Programmers usually don't give a darn about copyright other than being legally obligated to comply with it. So much of programming is copy-paste. GPL had its day, and it makes less and less sense as services take over. GPL locks small-time devs out of including the code in a for-profit project but does nothing to stop big corps from using it in a SaaS. The biggest irony is how Microsoft not only uses GPL'd code for profit but also ships WSL, all legally.
Re: Getty Images bans AI-generated content over fears of copyright claims
#217Earlier quoted context omitted.
A few seconds of a song is also very small, and yet it only take a few notes in some cases for a court to find someone guilty of copyright infringement. > On top of that, if the plaintiff could win, the actual market value of the source work is relevant to damages and that value is likely almost nothing so congratulations, you tied up the legal system just to get one AI-image taken down. The pirate bay trial gave pre…
Different domains of copyrightable material have different norms. The music industry, in response to sampling, has established that even small, recognizable snippets have marketable value and can therefore be infringed upon (there's it's also rife with case law with, in my opinion, bad wins by plaintiffs). For photographic images, collage is already an established art form and is generally considered transformative a…
It's my faint memory that for music, when it was just some kids looping breaks on vinyl printed in the hundreds or low thousands, the sampling was fine or at least not obviously an issue; but then copyright-holders saw that those artists and their management had started bringing in real money, and the laws were clarified.
Re: Getty Images bans AI-generated content over fears of copyright claims
#218Reading between the lines of this, it sounds to me like Getty is preparing a copyright claim against the AI companies: 1. They seem of the opinion that the copyright question is open. 2. Their business stands to lose substantially as a result of such models existing. 3. It would be a bad look for them to make a claim whilst simultaneously accepting works from the models into Getty. 4. At least some of their watermark…
Probably their value in the advertisment production chain is going to become close to zero in a few years, and they will try to stop or at least slow it down.
But once we have open source models released to the public I cannot see how a local legislation can have any impact at all.
Re: Getty Images bans AI-generated content over fears of copyright claims
#219Earlier quoted context omitted.
But it's not copyrightable. I guess you can lie and say you created it, but you didn't, and computer generated. You created it no more than you created your house because you picked the layout and paint colors. There is no money in non-copyrightable generated computer images.
Bullshit. I made a new image using a computer program that I was legally licensed to use. The program might be Corel Draw. It might be Blah Blah Diffusion Pro Plus. Either way, I made the image and I own the copyright, unless some other contract was made between myself and the program's owner or my employer.
If you put in creative inputs using a tool, it is copyrightable (a car's design). If all you did was say, give my XYZ widget (in this case 'give me a picture of a frog holding an umbrella under a rainbow') you only gave instructions for generating a widget, you did not create art.
Re: Getty Images bans AI-generated content over fears of copyright claims
#220Earlier quoted context omitted.
It’s not about reconstruction, it’s about the notion of a “derivative work”. Translating a work would absolutely be derivative (consider the case of translating a literary work between languages: this is a classic example of a derivative work). Blurring a work but incorporating it would nonetheless still be derivative, I think. The challenge with these models is that they’ve clearly been trained on (exposed to) copyr…
> If they were humans, a court could deem the outputs copyright infringement I'm not sure I understand how this is self-evident. The closest equivalent I can see would be a human who looks at many pieces of art to understand: - What is art and what is just scribbles or splatter? - What is good and what isn't? - What different styles are possible? Then the human goes and creates their own piece. It turns out, the lega…
The other challenges are: (i) the model isn't a human that can defend themself by explaining their creative process, it's a literal mathematical transformation of the inputs including the copyrighted work. (And I'm not sure "actually the human brain is just computation" defences offered by lawyers are ever likely to prevail in court, because if they do that opens much bigger cans of worms in virtually every legal field...) (ii) the representatives of OpenAI Inc who do have to explain themselves are going to have to talk about their approach to licenses for use of the material (which in this case appears to have been to disregard them altogether). That could be a serious issue for them even if the court agrees with the general principle that diffusion models or GANs are not plagiarism.
And possibly also (iii) the AI has ridiculous failure modes like implementing the Getty watermark which makes the model look far more closely derived from its source data than it actually is