Live data from Hacker News

Ask HN: DALL-E was trained on watermarked stock images?

news.ycombinator.com

81–90 of 233 posts

Re: Ask HN: DALL-E was trained on watermarked stock images?

#81
post #6

I remember when people used to say ianal. Innocent times when we thought there was an objective law and lawyers knew it. But that's not how these things work. The truth is that no one knows. Ultimately a bunch of people will decide how they feel about it. Well-read legal scholars trying really hard to be fair, but still just people. No one can predict with full certainty which way it will go.

> Ultimately a bunch of people will decide how they feel about it

Some would argue that technically these people _discover_[0] the law, but it amounts to the same thing

[0] https://www.jstor.org/stable/3143421

Re: Ask HN: DALL-E was trained on watermarked stock images?

#83

Kids in school are also trained on stock images https://www.reddit.com/r/KidsAreFuckingStupid/comments/8tgxs...

I think you're technically right, but that this will be overlooked from a legal perspective because it's less obvious that humans have been training ourselves on the prior art of others. We tend to blend in additional things besides prior art. (eg. nature, sensations, etc.)

Re: Ask HN: DALL-E was trained on watermarked stock images?

#85
Yes, Imagen and everything based on LAION 400M or 2B, too.

BTW, Copilot also ignored all licenses of the source code it memorized.

Datasets are the new capital. If they could, most employees would probably also object to their company using the result of their work to replace their job. But they can't. It's the same with artists here.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#86
post #48

Earlier quoted context omitted.

As I see it, 3 of the 4 tests are strongly in OpenAI's favor; the 'market effect' is mixed. (1) The use is highly transformative; (2) the images used were offered to the anonymous browsing public (with watermarks); (3) the end effect of training will only retain a tiny spectral distilled essence of any individual photo, or even a giant source corpus; (4) there's a potential risk of market competition from the ultimat…

These AI generated images are directly competing with stock images. AI tools are selling images to blogs and other customers that often would purchase stock images instead. The "character of use" is not in favor of dall-e, it is a commercial use. Copyright law does not require getty to block a user agents or ask them not to include their images. Another issue here is that removing copyright management info like a wat…

Whether something is directly competing for the same business would have to be evidenced, and copyright doesn't mean protection from all possible competition - it's just one factor weighed. And fair use protects many commercial uses, too, depending on proportion/character-of-original/etc.

But also, none of these images are direct, or even necessarily subtantial, "copies" of other images. The generator learned from other images – the same as any human artist might.

No watermark has been removed; the bigger issue may be that the spectral watermark violates a trademark. (But, I doubt consumers are likely to be confused.)

Re: Ask HN: DALL-E was trained on watermarked stock images?

#87

If you read the licence from Getty, they say, you are not allowed to use Getty pictures for ML.

What that license text says is irrelevant, because they’re not using it under that license, but under fair use exemptions in copyright law.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#90
post #74
post #48

Earlier quoted context omitted.

As I see it, 3 of the 4 tests are strongly in OpenAI's favor; the 'market effect' is mixed. (1) The use is highly transformative; (2) the images used were offered to the anonymous browsing public (with watermarks); (3) the end effect of training will only retain a tiny spectral distilled essence of any individual photo, or even a giant source corpus; (4) there's a potential risk of market competition from the ultimat…

I'm still not seeing the "transformative" argument: the point of transformation isn't "it is in a different format" but (to quote Wikipedia, which is, of course, dumb... I'm sorry ;P) where one "builds on a copyrighted work in a different manner or for a different purpose from the original". The reason a search engine thumbnail is transformative isn't because it has been transformed to make it smaller... it is becaus…

Autogenerated, often fantastical, never-seen-before AI images strike me as a paradigmatically 'transformative' use. It's novel. It's shocking to many practicioners how flexible & high-quality the images can be. It will unlock all sorts of new downstream creation.

The representation that feeds the generation is statistical, even to the point of being plausibly factual: these things/people/places/concepts can be abstractly represented as the balanced weights inside the model. And under US law, facts aren't copyrightable.

I could see a case being factored as: (1) the scraping/training/ephemeralization itself involves the usual copying of downloading/locally-processing images, like indexing, but all those 'copying' steps are fair-use protected, as science/transformative/de-minimus/whatever; (2) any subsequent new-image generation no longer involves any 'copying', only new creation from distilled patterns of the entire training corpus, in which Getty retains no 'trace tincture' of copyright-control. So there's no specific acts of illegal copying to penalize.

Also, a human artist would be allowed to review related Getty/etc preview images, free on the web, to familiarize themself with a person or setting, before drawing it themself, with their own flair – as long as they don't copy it substantially. Why wouldn't an AI artist?

Post reply on HN