Live data from Hacker News

Ask HN: DALL-E was trained on watermarked stock images?

news.ycombinator.com

141–150 of 233 posts

Re: Ask HN: DALL-E was trained on watermarked stock images?

#141
post #121

Earlier quoted context omitted.

Top 1% is a bit exaggerated, but there is definitely a lot of not good stuff. I find that Dall-E does especially poorly with underspecified prompts too, unlike something like Midjourney which can give visually pleasing photos for even the most abstract concepts. Dall-E tends to do better with concrete and specific prompts. Here's an example: Stressful Shapes Dall-E: https://i.imgur.com/JBkSh0y.png Midjourney: https:/…

But still "king of belgium giving a speech to an audience, but the audience members are cucumbers" is very specific. And I don't see the king of Belgium anywhere, two pictures have absolutely nothing to do with the prompt (no king, no speech, no audience, no cucumber), one has the speech and audience but no king or cucumber. Graphically, they are deep into the uncanny valley. Only the third image is kind of right, if…

IIRC, DALL-E filters requests related to politicians/celebrities. A friend had tried to make some funny stuff involving the Greek PM a couple of months ago, and it plainly refused. Now, it seems to process the request, but it will not show anyone resembling the person you asked for.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#144
post #133
post #30

I am not a lawyer, but I've had to argue about copyright with several. In the United States, there are two bits of case law that are widely cited and relevant: In Kelly v. Arriba Soft Corp (9th), found that making thumbnails of images for use in a search engine was sufficiently "transformative" that it was ok. Another case, Perfect 10 (9th), found that thumbnails for image search and cached pages were also transforma…

It seems it is possible to generate images which are very similar to the existing stock photos if you feed getty images' description into DALL-E. I tried it with a distinctive banana image: https://imgur.com/a/0OrIr6e

Interesting. Adding "stock photo" to the string generated that getty tag? That is probably the most attackable (alas easy to fix) part of the issue. It will be an interesting question how close to the original a picture has to be to be considered the same (I'm sure there's some case law) and maybe there's some new research to be done regarding how to recreate the training data images with the correct search string (I suppose one could build an ML model for that).

Fun times ahead

Re: Ask HN: DALL-E was trained on watermarked stock images?

#146

Earlier quoted context omitted.

I have been saying for a long time. These big companies with an investment in getting these systems working could collaborate with twitter and instagram to include image license options. They could give users the option to opt-in to data collection. This would help the public feel like they’re giving consent, and it would also lead to the creation of massive, properly licensed open data sets. There’s already some big…

> They could give users the option to opt-in to data collection. This would help the public feel like they’re giving consent Would we get paid? I'd care a lot less about OpenAI profiting from my work if I was getting commission every time my work was hit in the training data.

I guess this is what those NFT supporters were talking about.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#148
post #86

Earlier quoted context omitted.

These AI generated images are directly competing with stock images. AI tools are selling images to blogs and other customers that often would purchase stock images instead. The "character of use" is not in favor of dall-e, it is a commercial use. Copyright law does not require getty to block a user agents or ask them not to include their images. Another issue here is that removing copyright management info like a wat…

Whether something is directly competing for the same business would have to be evidenced, and copyright doesn't mean protection from all possible competition - it's just one factor weighed. And fair use protects many commercial uses, too, depending on proportion/character-of-original/etc. But also, none of these images are direct, or even necessarily subtantial, "copies" of other images. The generator learned from ot…

It’s going to be interesting what the stock companies will do. Maybe they will make their own Image Generator. Perhaps we will see a case based on the new factor that is AI. An AI is not artist; they can’t be conflated. A decent artists can churn out maybe 5-10 works if he is productive. AI can churn out by the hundreds or thousands if needed. The process also isn’t the same.

Anyway it will be interesting to watch this space.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#149
post #144
post #133

Earlier quoted context omitted.

It seems it is possible to generate images which are very similar to the existing stock photos if you feed getty images' description into DALL-E. I tried it with a distinctive banana image: https://imgur.com/a/0OrIr6e

Interesting. Adding "stock photo" to the string generated that getty tag? That is probably the most attackable (alas easy to fix) part of the issue. It will be an interesting question how close to the original a picture has to be to be considered the same (I'm sure there's some case law) and maybe there's some new research to be done regarding how to recreate the training data images with the correct search string (I…

No, I didn't get the tag. But I suppose that Getty metadata as well as the images were used for training.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#150
post #48

Earlier quoted context omitted.

Search engines don't create market harm for a work because they don't compete with it. In fact, they do the opposite: they advertise the work, making it more accessible and increasing exposure. These AI tools on the other hand seem to do the exact opposite. They can (or could, if they got good enough) absolutely compete with a work, and therefore seem like they create substantial market harm. The character of use als…

As I see it, 3 of the 4 tests are strongly in OpenAI's favor; the 'market effect' is mixed. (1) The use is highly transformative; (2) the images used were offered to the anonymous browsing public (with watermarks); (3) the end effect of training will only retain a tiny spectral distilled essence of any individual photo, or even a giant source corpus; (4) there's a potential risk of market competition from the ultimat…

I have a hard time agreeing with 3, given https://ibb.co/DzGR063
Post reply on HN