Live data from Hacker News

Getty Images bans AI-generated content over fears of copyright claims

theverge.com

161–170 of 390 posts

Re: Getty Images bans AI-generated content over fears of copyright claims

#161
post #147

It's always weird to see the contrast between HN's reaction to copyright questions about text/image generation, and HN's reaction when it's code generation. When a model is trained on 'all-rights-reserved' content like most image datasets, the community say it's fair game. But when it's 'just-a-few-rights-reserved' content like GPL code, apparently the community says that crosses a line? Realistically, this tells me…

If this was a website for artists instead of programmers you'd see the exact opposite pattern. Unfortunately, people only seem to care when it threatens their own livelihood, not when it threatens that of the people around them.

The average HN denizen gets so pissed off when I ask them to save their comment rejoicing in the inevitability of the elimination of my entire field, and take it out when next year's descendant of CodePilot gives them the same horrible sinking feeling that these things give me.

Re: Getty Images bans AI-generated content over fears of copyright claims

#162

It is very easy to see hypocrisy in the FurAffinty statement: "Human artists see, analyze and even sample other artists’ work to create content. That content generated can reference hundreds, even thousands of pieces of work from other artists that they have consumed in their lifetime to create derivative images,” ... “Our goal is to support artists and their content. We don’t believe it’s in our community’s best int…

generated images =/= created AI generated images have the potential to be art in the eyes of the beholder, but let's not pretend that generation is the same as the mental, physical, and spiritual flow state that goes into painting or drawing a piece.

I don't think that's what the parent is doing. They're pointing out the hypocrisy of claiming AI art is copying copyrighted works because human artists are trained in similar ways. That's not making a claim about whether or not AI art is "real" art.

Re: Getty Images bans AI-generated content over fears of copyright claims

#163

Earlier quoted context omitted.

I've seen a lot of confidence on HN and other tech communities that a court would never rule that training an AI on copyrighted images is infringement, but I'm not so sure. To be clear, I hope that training AI on copyrighted images remains legal, because it would cripple the field of AI text and image generation if it wasn't! But think about these similar hypotheticals: 1. I take a copyrighted Getty stock image (that…

It’s not about reconstruction, it’s about the notion of a “derivative work”. Translating a work would absolutely be derivative (consider the case of translating a literary work between languages: this is a classic example of a derivative work). Blurring a work but incorporating it would nonetheless still be derivative, I think. The challenge with these models is that they’ve clearly been trained on (exposed to) copyr…

> If they were humans, a court could deem the outputs copyright infringement

I'm not sure I understand how this is self-evident. The closest equivalent I can see would be a human who looks at many pieces of art to understand:

- What is art and what is just scribbles or splatter?

- What is good and what isn't?

- What different styles are possible?

Then the human goes and creates their own piece.

It turns out, the legal solution is to evaluate each piece individually rather than the process. And, within that, the court has settled on "if it looks like a duck and it quacks like a duck..." which is where the subconscious copying presumably comes in.

I don't know where courts will go. The new challenge is AI can generate "potentially infringing" work at a much higher rate than humans, but that's really about it. I'd be surprised if it gets treated materially different than human-created works.

Re: Getty Images bans AI-generated content over fears of copyright claims

#164
post #154

Earlier quoted context omitted.

> if I copy and paste one page from each of a thousand books it would be. It almost certainly would not be infringement. One page of text out of an entire work is very small. Amount and substantiality of the portion used in relation to the copyrighted work as a whole is one of the factors considered when making a fair use defense. This hypothetical book would also have zero effect on the potential market for the sour…

A few seconds of a song is also very small, and yet it only take a few notes in some cases for a court to find someone guilty of copyright infringement. > On top of that, if the plaintiff could win, the actual market value of the source work is relevant to damages and that value is likely almost nothing so congratulations, you tied up the legal system just to get one AI-image taken down. The pirate bay trial gave pre…

Different domains of copyrightable material have different norms. The music industry, in response to sampling, has established that even small, recognizable snippets have marketable value and can therefore be infringed upon (there's it's also rife with case law with, in my opinion, bad wins by plaintiffs). For photographic images, collage is already an established art form and is generally considered transformative and fair use of the source images.

I'm not familiar with any The Pirate Bay case; if you are referring to this one [0], it was in a Swedish court and I'm not familiar with Swedish copyright law. However, the first sentence says the charge was promoting infringement, not that they were engaged in infringement themselves. I don't think that's relevant to what I was replying to but could be very relevant to Getty Images's decision, if AI-generate content is infringing, they don't want to be accused of promoting infringement. There's undoubtedly already infringement taking place on Getty Images but likely at such a small scale that the organization itself is not put at risk.

[0] https://en.wikipedia.org/wiki/The_Pirate_Bay_trial

Re: Getty Images bans AI-generated content over fears of copyright claims

#165

Earlier quoted context omitted.

It’s not about reconstruction, it’s about the notion of a “derivative work”. Translating a work would absolutely be derivative (consider the case of translating a literary work between languages: this is a classic example of a derivative work). Blurring a work but incorporating it would nonetheless still be derivative, I think. The challenge with these models is that they’ve clearly been trained on (exposed to) copyr…

> If they were humans, a court could deem the outputs copyright infringement I'm not sure I understand how this is self-evident. The closest equivalent I can see would be a human who looks at many pieces of art to understand: - What is art and what is just scribbles or splatter? - What is good and what isn't? - What different styles are possible? Then the human goes and creates their own piece. It turns out, the lega…

It's worth pointing out that the problem, in this scenario, is for the creator (ie, the human running the algorithm). They will need to determine whether a piece might violate copyright before using it or selling it. That seems like a very hard problem, and could be the justification for more [new] blanket rules on the AI process.

Re: Getty Images bans AI-generated content over fears of copyright claims

#166

Earlier quoted context omitted.

https://www.smithsonianmag.com/smart-news/us-copyright-offic... > An image generated through artificial intelligence lacked the “human authorship” necessary for protection > Both in its 2019 decision and its decision this February, the USCO found the “human authorship” element was lacking and was wholly necessary to obtain a copyright, Engadget’s K. Holt wrote. Current copyright law only provides protections to “the…

Take a look at the decision: https://www.copyright.gov/rulings-filings/review-board/docs/... The person was still trying to get the AI marked as the owner. In fact in the application "does not assert that the Work was created with contribution from a human author" which the office acceded to but did not actually agree or disagree with. So it still says nothing about whether a human can have copyright over a image the…

A distinction without a difference, since this was just someone who was hoping to be the beneficial owner of an AI with an enforceable copyright interest. Recall the failure of the photographer who allowed monkeys to play with his camera equipment, leading one of them to take a selfie photo that became famous.

The photographer asserted copyright on the basis that he had brought his camera there, befriended the monkeys, and set his equipment up in such a way that even a monkey could use it and get a quality image, but his claim to authorship was rejected and so he was unable to realize any profit from selling the photo - although I'm sure he made it up on speaking tours telling the story of how he got it.

To be sure, AI created art is done in response to a prompt provided by a human, but unless that human has done all the training and calculation of weights, they can't claim full ownership on the output from the model. There's a stronger case where a human supplies an image prompt and the textual input describes stylistic rather than structural content.

Re: Getty Images bans AI-generated content over fears of copyright claims

#167

Earlier quoted context omitted.

https://www.smithsonianmag.com/smart-news/us-copyright-offic... > An image generated through artificial intelligence lacked the “human authorship” necessary for protection > Both in its 2019 decision and its decision this February, the USCO found the “human authorship” element was lacking and was wholly necessary to obtain a copyright, Engadget’s K. Holt wrote. Current copyright law only provides protections to “the…

Doesn't seem to apply to Stable Diffusion and Dall-E because there is substantial human work involved - picking and evolving the prompt and selecting the best result. Sometimes it's also collaging and masking. Maybe it could apply to making "variations" where you just have to click a button. But you still have to choose the subject image on which you do variations, and to pick the best one or scrap the lot. It's comp…

> a monkey stealing your camera

It's really not different, because while everyone remembers it that way, the photographer went to great lengths to facilitate the monkey selfie.

https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...

Re: Getty Images bans AI-generated content over fears of copyright claims

#168
post #68
post #62

Reading between the lines of this, it sounds to me like Getty is preparing a copyright claim against the AI companies: 1. They seem of the opinion that the copyright question is open. 2. Their business stands to lose substantially as a result of such models existing. 3. It would be a bad look for them to make a claim whilst simultaneously accepting works from the models into Getty. 4. At least some of their watermark…

> 1. They seem of the opinion that the copyright question is open. I'm surprised it has taken this long to be honest. I've seen generated images with the blurred Getty watermark on them.

It's not that the watermark is on them per se, but that the model tried to emulate an image it had seen before which had a watermark on it. Imagine showing a child a bunch of pictures with Getty watermarks on them, then they draw their own, with their own emulation of the watermark. They don't know it's a watermark, they don't know what a watermark is, they just see this shape on a lot of pictures and put it on their own. That's essentially what's going on.

The model is only around 4GB, and it was trained on ~5B images. At 24 bit depth, that'd be 786k raw data per image, which would be 3.5 petabytes of uncompressed information in the full training set. Either the authors have invented the world's greatest compression algorithm, or the original image data isn't actually in the model.

So, I think the argument is: if you look at someone else's (copyrighted) work, and produce your own work incorporating style, composition, etc elements which you learned from their work, are you engaged in copyright infringement? IANAL but I think the answer is "no" - you would have to try to reproduce the actual work to be engaging in copyright infringement, and not only do these models not do that, it would be extremely hard to get them to do so without feeding them the actual copyrighted work as an input to the inference procedure.

Re: Getty Images bans AI-generated content over fears of copyright claims

#169

Earlier quoted context omitted.

The US Copyright Office asserts that AI generated images can’t be copyrighted. Getty lives and dies by copyright and artificial scarcity/control of image rights. For stock images and non current/news events, Stable Diffuison and its successors are the future.

You still need computing power to generate those images, so definitely room for commercial activity there. Getty could precompute billions of images and enlarge their inventory.

Sure, bt why should anyone waste time browsing their inventory when they can just make up their own?

Re: Getty Images bans AI-generated content over fears of copyright claims

#170

It's always weird to see the contrast between HN's reaction to copyright questions about text/image generation, and HN's reaction when it's code generation. When a model is trained on 'all-rights-reserved' content like most image datasets, the community say it's fair game. But when it's 'just-a-few-rights-reserved' content like GPL code, apparently the community says that crosses a line? Realistically, this tells me…

that's bc hn has a lot of gpl zealots who want special rules because they believe a viral license is a "fundamental good" and closing stuff off via copyright is the opposite. i don't agree and think viral licenses suck, but it's not a unpopular opinion on here.
Post reply on HN