Regardless of whether or not training an AI on stock images violates the license, there's a very real problem with that watermark being present, which is that it proves their AI is prone to copying large swaths of images from gettyimages unaltered, and that definitely is a license violation. This makes me think back to the controversy over github copilot; if these AIs are going to be trained on other peoples' IP then…
Ask HN: DALL-E was trained on watermarked stock images?
91–100 of 233 posts
Re: Ask HN: DALL-E was trained on watermarked stock images?
#92All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…
To me this feels like the argument that we should allow Uber and Airbnb because they're sufficiently "transformative" use cases. When clearly they are playing fast and loose by the rules and have taken advantage of being early enough to do so. As soon as the rulemakers caught up, it became obvious that they didn't have a license to operate differently from everyone else, just because they're new and popular. Personal…
This is exactly right. Open Datasets is the way to go. I would also say that in the spirit of the Open Access movement for journals and publications, it might be useful to set up an Open Access protocol for training data sets, methods (these are just the algorithms; publishing them openly might be the way to go) and computed models.
This will ensure that models are evaluated for risk by a large set of people and any risks/shortcomings could be addressed soon. Quite similar to how cryptographic algorithms are designed / analyzed in public. Obfuscation might look like it helps, but it doesn't in the long run and just creates more headaches.
Re: Ask HN: DALL-E was trained on watermarked stock images?
#93Earlier quoted context omitted.
yeah except this artist wont go around painting watermarks
Really? https://www.moma.org/learn/moma_learning/andy-warhol-campbel... (Yes I get it's not technically a watermark, but it certainly qualifies as a trade mark in a similar fashion)
Re: Ask HN: DALL-E was trained on watermarked stock images?
#94Earlier quoted context omitted.
Of course people are more likely to share the best iamges – or in this case, the one most illustrative of their concern (about watermarks). Also: my sense is that getting the best results often requires a lot of extra coaching with style/detail words. As we can't see the prompt here, we don't know what sort of style/details were requested. GIGO.
OP did say what prompt they used
There's one screenshot showing the prompt – but in $CURRENT_YEAR, I view all screenshots with at least a little suspicion, especially when there was a way to highlight the pseuod-watermarked image – OpenAI's native 'Share' – that would've provided stronger proof, direct from OpenAI, of exactly the prompt associated with an image. Hoaxes are everywhere! I've added DALL-E bottom-right color-squares to non-DALL-E images, & seen others do the same, as a subtle joke.
So I generally believe the OP, but don't rule-out the possibility there's been tampering to make some point.
Re: Ask HN: DALL-E was trained on watermarked stock images?
#95Re: Ask HN: DALL-E was trained on watermarked stock images?
#96These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.
Re: Ask HN: DALL-E was trained on watermarked stock images?
#97They probably already have specialized filtering models built to filter out censorable terms. They may be imperfect, but they are there. A watermark remover might be an easy addition.
When Stable Diffusion released their model playground, I used the prompt Peter at the pearly gates dressed as a security guard and got three images two of which were censored and one that was an ordinary image. So, the capability is there already. Just a matter of time before they get good at watermark removal.
Re: Ask HN: DALL-E was trained on watermarked stock images?
#98These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.
On top of that DALL-E2 has generally issues with anything dealing with multiple objects. A single person will render fine, groups of people will generally give artifacts. Attributes will also be spread across all objects in the scenes, not just the ones you specified in your prompt, so doing anything more complex will require manual uncropping und inpainting, not just a single prompt.
Anyway, if you avoid the obvious weak spots and holes in the training set, DALL-E2 output is for most part pretty amazing out of the box. It's really more a top 50% than a top 1%.
The biggest bias when it comes to published DALL-E2 images are the prompts. Most prompts you see online are not the actual prompts, but funny descriptions made by a human after the fact. The actual prompt are often much longer and sometimes completely different.
Re: Ask HN: DALL-E was trained on watermarked stock images?
#99These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.
Let's not confuse the AI with "buts", just say that he is giving the speech to cucumbers.
Lastly, specify some style, because this would probably not work out as a photo.
My single try is not bad at all and it could definitely be improved.
Re: Ask HN: DALL-E was trained on watermarked stock images?
#100These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.
Yes, people tend to share the best of the best. However these results seem especially bad, like bottom 10% bad.