Live data from Hacker News

Ask HN: DALL-E was trained on watermarked stock images?

news.ycombinator.com

91–100 of 233 posts

Re: Ask HN: DALL-E was trained on watermarked stock images?

#91

Regardless of whether or not training an AI on stock images violates the license, there's a very real problem with that watermark being present, which is that it proves their AI is prone to copying large swaths of images from gettyimages unaltered, and that definitely is a license violation. This makes me think back to the controversy over github copilot; if these AIs are going to be trained on other peoples' IP then…

Just because it contains the text of the watermark does not mean that it's reproducing large swaths of the image - its doubtful even the most generous perceptive hash would retrieve any matching images in the Getty repository.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#92

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

To me this feels like the argument that we should allow Uber and Airbnb because they're sufficiently "transformative" use cases. When clearly they are playing fast and loose by the rules and have taken advantage of being early enough to do so. As soon as the rulemakers caught up, it became obvious that they didn't have a license to operate differently from everyone else, just because they're new and popular. Personal…

> What it would do is to create a market for open source data sets with liberal licenses.

This is exactly right. Open Datasets is the way to go. I would also say that in the spirit of the Open Access movement for journals and publications, it might be useful to set up an Open Access protocol for training data sets, methods (these are just the algorithms; publishing them openly might be the way to go) and computed models.

This will ensure that models are evaluated for risk by a large set of people and any risks/shortcomings could be addressed soon. Quite similar to how cryptographic algorithms are designed / analyzed in public. Obfuscation might look like it helps, but it doesn't in the long run and just creates more headaches.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#93
post #26

Earlier quoted context omitted.

yeah except this artist wont go around painting watermarks

Really? https://www.moma.org/learn/moma_learning/andy-warhol-campbel... (Yes I get it's not technically a watermark, but it certainly qualifies as a trade mark in a similar fashion)

Ceci n'est pas une pipe … er … soupe

Re: Ask HN: DALL-E was trained on watermarked stock images?

#94
post #60
post #41

Earlier quoted context omitted.

Of course people are more likely to share the best iamges – or in this case, the one most illustrative of their concern (about watermarks). Also: my sense is that getting the best results often requires a lot of extra coaching with style/detail words. As we can't see the prompt here, we don't know what sort of style/details were requested. GIGO.

OP did say what prompt they used

I've often seen people show off their autogenerated images and report only approximate paraphrases of their actual prompts.

There's one screenshot showing the prompt – but in $CURRENT_YEAR, I view all screenshots with at least a little suspicion, especially when there was a way to highlight the pseuod-watermarked image – OpenAI's native 'Share' – that would've provided stronger proof, direct from OpenAI, of exactly the prompt associated with an image. Hoaxes are everywhere! I've added DALL-E bottom-right color-squares to non-DALL-E images, & seen others do the same, as a subtle joke.

So I generally believe the OP, but don't rule-out the possibility there's been tampering to make some point.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#96

These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.

No post body was provided.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#97
Just wait until they build an AI watermark identifier and remover (which is a problem subset) and then use its output to train/update their model.

They probably already have specialized filtering models built to filter out censorable terms. They may be imperfect, but they are there. A watermark remover might be an easy addition.

When Stable Diffusion released their model playground, I used the prompt Peter at the pearly gates dressed as a security guard and got three images two of which were censored and one that was an ordinary image. So, the capability is there already. Just a matter of time before they get good at watermark removal.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#98

These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.

Um, have you read the prompt? It looking weird is simply the result of "the audience members are cucumbers". The more crazy your prompt is, the worse the results will generally get.

On top of that DALL-E2 has generally issues with anything dealing with multiple objects. A single person will render fine, groups of people will generally give artifacts. Attributes will also be spread across all objects in the scenes, not just the ones you specified in your prompt, so doing anything more complex will require manual uncropping und inpainting, not just a single prompt.

Anyway, if you avoid the obvious weak spots and holes in the training set, DALL-E2 output is for most part pretty amazing out of the box. It's really more a top 50% than a top 1%.

The biggest bias when it comes to published DALL-E2 images are the prompts. Most prompts you see online are not the actual prompts, but funny descriptions made by a human after the fact. The actual prompt are often much longer and sometimes completely different.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#99

These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.

Op constructed a horrible prompt. First of all, using king Philippe I. is against the ToS, so let's go with a generic "king".

Let's not confuse the AI with "buts", just say that he is giving the speech to cucumbers.

Lastly, specify some style, because this would probably not work out as a photo.

My single try is not bad at all and it could definitely be improved.

https://labs.openai.com/s/3OUmUxKefJCeLhAk4hkeKX4V

Re: Ask HN: DALL-E was trained on watermarked stock images?

#100

These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.

They are the worst I’ve seen as well.

Yes, people tend to share the best of the best. However these results seem especially bad, like bottom 10% bad.

Post reply on HN