Live data from Hacker News

Getty Images bans AI-generated content over fears of copyright claims

theverge.com

271–280 of 390 posts

Re: Getty Images bans AI-generated content over fears of copyright claims

#271

Earlier quoted context omitted.

Stable Diffusion is already out in the world. The cat is out of the bag.

Yah, but think of Napster getting eventually usurped by Spotify. The danger is that it's legally no longer possible to update the models (which are very expensive to train), and we end up with only Disney with the copyright horde large enough to train decent models, let alone good ones...

The models are expensive to train right now, but I suspect in 10 years, anyone with a multi gpu rig could train the equivalent of Stable Diffusion.

Re: Getty Images bans AI-generated content over fears of copyright claims

#272
post #62

Reading between the lines of this, it sounds to me like Getty is preparing a copyright claim against the AI companies: 1. They seem of the opinion that the copyright question is open. 2. Their business stands to lose substantially as a result of such models existing. 3. It would be a bad look for them to make a claim whilst simultaneously accepting works from the models into Getty. 4. At least some of their watermark…

I've seen a lot of confidence on HN and other tech communities that a court would never rule that training an AI on copyrighted images is infringement, but I'm not so sure. To be clear, I hope that training AI on copyrighted images remains legal, because it would cripple the field of AI text and image generation if it wasn't! But think about these similar hypotheticals: 1. I take a copyrighted Getty stock image (that…

There's a literal arms race behind the scenes in AI right now.

I think it's very unlikely that corporate IP claims that could substantially hold back progress in domestic AI development will end up being successful.

Re: Getty Images bans AI-generated content over fears of copyright claims

#273
post #213

Earlier quoted context omitted.

That's not so obviously clear cut. Can the model produce identical output from the same simple prompt?

But that's exactly what happens, AI isn't randomness, it's a set of predefined calculations. The randomness is in the seed/starting point. For e.g Stable Diffusion it is given/user input, resulting in perfect reproducibility.

I should have been clearer - that was rhetorical to point out that if you and I use the same prompt and pick the same resultant image, it's harder to claim either of us have copyright. This is a gray area. The tools themselves could introduce some stochastic aspect so that outputs are never identical, also.

None of this leads to an obviously clear cut legal position wrt copyright.

Re: Getty Images bans AI-generated content over fears of copyright claims

#274

It's always weird to see the contrast between HN's reaction to copyright questions about text/image generation, and HN's reaction when it's code generation. When a model is trained on 'all-rights-reserved' content like most image datasets, the community say it's fair game. But when it's 'just-a-few-rights-reserved' content like GPL code, apparently the community says that crosses a line? Realistically, this tells me…

When a model is trained on 'all-rights-reserved' content like most image datasets, the community say it's fair game. But when it's 'just-a-few-rights-reserved' content like GPL code, apparently the community says that crosses a line? A) This is just taking divided opinion and treating it like a person with a contradictory opinion (as others have noted). B) Nothing about GPL makes it "less copyrighted". Acting like a…

> Nothing about GPL makes it "less copyrighted". Acting like a commercial copyright is "stronger" because it doesn't immediately grant certain uses is false and needs to be challenged whenever the claim is made.

GPL says "you have a license to use it if you do XYZ." The alternative is "you have no license to use it." How is that not strictly "stronger?"

Re: Getty Images bans AI-generated content over fears of copyright claims

#275

It's always weird to see the contrast between HN's reaction to copyright questions about text/image generation, and HN's reaction when it's code generation. When a model is trained on 'all-rights-reserved' content like most image datasets, the community say it's fair game. But when it's 'just-a-few-rights-reserved' content like GPL code, apparently the community says that crosses a line? Realistically, this tells me…

> When a model is trained on 'all-rights-reserved' content like most image datasets, the community say it's fair game. But when it's 'just-a-few-rights-reserved' content like GPL code, apparently the community says that crosses a line? I don't think this is right. I think different people have different views, and you're just assuming that the same people have contradictory views.

I'm not assuming that, but I see how it could read that way, since I'm being fast and loose with the language. The community (anthropomorphizing the blob again) definitely empirically reacts very differently to the two topics.

Re: Getty Images bans AI-generated content over fears of copyright claims

#276

It's always weird to see the contrast between HN's reaction to copyright questions about text/image generation, and HN's reaction when it's code generation. When a model is trained on 'all-rights-reserved' content like most image datasets, the community say it's fair game. But when it's 'just-a-few-rights-reserved' content like GPL code, apparently the community says that crosses a line? Realistically, this tells me…

People don't like being told no. The vast majority of all-rights-reserved content is either not licensable, or not licensable at a price that anyone would be willing to pay or can afford. Ergo we[0] would much rather see more opportunities to use the work without needing permission, because we will never have permission . When getting permission is reasonable then people are willing to defend the system. And code is…

That's a good way to look at it: The difference between "no" and "yes if XYZ."

Maybe it's that people respect "yes if XYZ" more than "no" because there's some path to yes that way. In that case, it really does speak to the need for some open-ish text and image licenses, like "you can use my image in a model if you share the model with me."

Re: Getty Images bans AI-generated content over fears of copyright claims

#277

It's always weird to see the contrast between HN's reaction to copyright questions about text/image generation, and HN's reaction when it's code generation. When a model is trained on 'all-rights-reserved' content like most image datasets, the community say it's fair game. But when it's 'just-a-few-rights-reserved' content like GPL code, apparently the community says that crosses a line? Realistically, this tells me…

> You could imagine a GPL-like license being really good for the community/ecosystem: "If you train on this content, you have to release the model." The issue with that is that it is perfectly legal to ignore the license, if you use the copyrighted work in a transformative way. It doesn't matter what the license says, if it is legal to ignore the license.

That seems to be the status quo, but if huge companies benefit from the electorate's work and use that to put them out of jobs, I wouldn't be surprised if the law changes.

Re: Getty Images bans AI-generated content over fears of copyright claims

#278
post #240

Earlier quoted context omitted.

I think you'd struggle to argue that the Getty watermark was a general style and composition principle and not a distinct motif unique to Getty (and in music copyright cases, the defence of plagiarising motifs inadvertently frequently fails).

From the model's perspective, it's not a distinct motif, that's the thing (and, it struggles quite a lot to reproduce the actual mark). The model doesn't have any concept of what a "watermark" is. As far as it's concerned, it's just a compositional element that happens to be in some images. Most "watermarks" Stable Diffusion produces are jumbles of colorized pixels which we can recognize as being evocative of a water…

> From the model's perspective, it's not a distinct motif, that's the thing (and, it struggles quite a lot to reproduce the actual mark). The model doesn't have any concept of what a "watermark" is.

The court delivers the judgement, not the model.

If courts can find against musicians whilst accepting they 'unconsciously' plagiarised key elements of a song in their own completely different song played by different musicians based on maybe hearing it in the background somewhere, they can certainly find against the creators of a model which has a sufficiently strong and obvious dependency on Getty IP they imported to output reasonably close approximations of Getty watermarks.

Re: Getty Images bans AI-generated content over fears of copyright claims

#279

Earlier quoted context omitted.

I believe the more important issue is that material generated through AI is not copyrightable by the author of such images. If you are an artist you can't claim any copyright on what you're generating. If you're not the copyrighted holder it follows that you can't sell it or that you can't complain if someone else copy it (verbatim) and sells it.

The criteria for a work being copyrightable literally is the slightest touch of creativity, and I'm quite certain that writing a prompt and selecting a result out of a bunch of random seeds would qualify for that. Fully automated mass creation would get excluded, as would be any attempts to assert that copyright to a non-human entity, but all the artwork I've seen generated by people should be copyrightable - the mai…

> The criteria for a work being copyrightable literally is the slightest touch of creativity

If a prompt is copyrightable, that's a problem. Because it's just words. Recipes should be in the same league then.

If I can get the same output with a slightly different prompt, how would you protect your works?

If I copy your output, how can you protect your works, given that the output depends on something not copyrightable? (as per your statement, which I agree with)

Look at it this way: If I make something out of a Spirograph, is it a copyrightable work?

Re: Getty Images bans AI-generated content over fears of copyright claims

#280
post #265
post #245

Earlier quoted context omitted.

Yes... ish. On one hand, if you do a "this is the size of the net" and then divide it by the number of training images, its rather small amount of storage per image. On the other hand, when I was playing with stable diffusion on the command line following the instructions of https://replicate.com/blog/run-stable-diffusion-on-m1-mac python scripts/txt2img.py --prompt "wolf with bling walking down a street" --n_samples…

That's a very interesting result. Did you happen to capture the seed for either of those first two images? It would be interesting to try to reproduce.

Alas no. And I haven't been able to tickle it again in the right way to get those images out.

The invocation of that run is still in my scroll back:

    (venv) shagie@MacM1 stable-diffusion % python scripts/txt2img.py --prompt "wolf with bling walking down a street" --n_samples 6 --n_iter 1 --plms
    Global seed set to 42
    Loading model from models/ldm/stable-diffusion-v1/model.ckpt
    Global Step: 470000
    LatentDiffusion: Running in eps-prediction mode
    DiffusionWrapper has 859.52 M params.
    making attention of type 'vanilla' with 512 in_channels
    Working with z of shape (1, 4, 32, 32) = 4096 dimensions.
    making attention of type 'vanilla' with 512 in_channels
That's the only spot I see the seed mentioned and then it goes on with lots of other logging but nothing seed related that would indicate a way to reproduce it.

---

(late edit) you can fairly accurately (so far 1 image out of 20) get that image out with the prompt "Rick Astley Never Gonna Give You Up"

Post reply on HN