Earlier quoted context omitted.
Another human analogy could be: you take a photo from every art show and museum, and use those for reference as you paint.
The analogy of using your cortex is more apt. My understanding of neural networks is that there are no remnants of the original inside it. The training data is used to back propagate a bunch of weights. Your brain works like those neural network neurons; they learn when to fire, but they don’t know the intricate detail like a photo. Hence why many claim eyewitness testimony is bogus.
Ask HN: DALL-E was trained on watermarked stock images?
101–110 of 233 posts
Re: Ask HN: DALL-E was trained on watermarked stock images?
#102> but surely you can't just... use stock photos without paying for the license? They aren't hosting the infringing content. Training on the data is probably covered under fair use. Generations are of _learned_ representations of the dataset, not the dataset itself. This makes it closer to outputting original works (probably owned by the person who used the model). The players involved here are known for being litigio…
> Training on the data is probably covered under fair use. Generations are of _learned_ representations of the dataset, not the dataset itself. This makes it closer to outputting original works (probably owned by the person who used the model). "Probably" is doing a lot of heavy lifting in that sentence. As for "_learned_", that's pretty debatable considering it's reproducing recognizable trademark infringement. > Th…
Copying ideas and styles has always been a fundamental part of art history, so an artwork right holder might have a hard time successfuly sueing a user for the user's generated image looking similar to the right holder's artwork.
Re: Ask HN: DALL-E was trained on watermarked stock images?
#103These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.
Um, have you read the prompt? It looking weird is simply the result of "the audience members are cucumbers". The more crazy your prompt is, the worse the results will generally get. On top of that DALL-E2 has generally issues with anything dealing with multiple objects. A single person will render fine, groups of people will generally give artifacts. Attributes will also be spread across all objects in the scenes, no…
Perhaps this rewrite may yield better results:
"King of Belgium gives a speech to an audience of cucumbers"
Re: Ask HN: DALL-E was trained on watermarked stock images?
#104These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.
Re: Ask HN: DALL-E was trained on watermarked stock images?
#105There are lots of photos with watermark circulating on web, for example in memes and unfinished webpages (when finished, these will be replaced with paid variant without watermark).
Re: Ask HN: DALL-E was trained on watermarked stock images?
#106These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.
Top 1% is a bit exaggerated, but there is definitely a lot of not good stuff. I find that Dall-E does especially poorly with underspecified prompts too, unlike something like Midjourney which can give visually pleasing photos for even the most abstract concepts. Dall-E tends to do better with concrete and specific prompts. Here's an example: Stressful Shapes Dall-E: https://i.imgur.com/JBkSh0y.png Midjourney: https:/…
What puzzles me is if the Getty Images logo can sometimes appear. If you only have a Getty account, you get rid of the logo and can legally use them royalty free?
Re: Ask HN: DALL-E was trained on watermarked stock images?
#107> but surely you can't just... use stock photos without paying for the license? They aren't hosting the infringing content. Training on the data is probably covered under fair use. Generations are of _learned_ representations of the dataset, not the dataset itself. This makes it closer to outputting original works (probably owned by the person who used the model). The players involved here are known for being litigio…
> Training on the data is probably covered under fair use. Generations are of _learned_ representations of the dataset, not the dataset itself. This makes it closer to outputting original works (probably owned by the person who used the model). "Probably" is doing a lot of heavy lifting in that sentence. As for "_learned_", that's pretty debatable considering it's reproducing recognizable trademark infringement. > Th…
Indeed, that's why I used it. It wasn't long ago that DALLE-2 outputs were the ownership of OpenAI (they changed it so the owner is the user recently). Definitely plenty of room for debate on who the owner should be.
> As for "_learned_", that's pretty debatable considering it's reproducing recognizable trademark infringement.
I guess. I meant this strictly in the machine learning sense, where "learned" is typically used to describe models trained via stochastic gradient descent.
> I have no idea why anyone would assume the "move fast and break things" disruption mindset that pervades tech companies these days, especially in spaces like ML/"AI", would mean they considered the legality, ethics, or good business sense of their training dataset.
I agree mostly, except that companies like Alamy have their hooks in everywhere so they can seek rent. I just figured they might be cautious about this if e.g. Microsoft (OpenAI's business partner) had an existing agreement in place for Bing or something.
Re: Ask HN: DALL-E was trained on watermarked stock images?
#108All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…
My feeling is that while these things may be technically falling under fair use, I really feel like they are running roughshod over a lot of ethical and moral lines and that perhaps "fair use" needs to be redefined to explicitly exclude this kind of processing. And if it kills these things, oh well. "Being an artist" is a precarious enough existence in this world as is, I'd be delighted to stop worrying about having…
Re: Ask HN: DALL-E was trained on watermarked stock images?
#109I am not a lawyer, but I've had to argue about copyright with several. In the United States, there are two bits of case law that are widely cited and relevant: In Kelly v. Arriba Soft Corp (9th), found that making thumbnails of images for use in a search engine was sufficiently "transformative" that it was ok. Another case, Perfect 10 (9th), found that thumbnails for image search and cached pages were also transforma…
Re: Ask HN: DALL-E was trained on watermarked stock images?
#110These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.
Op constructed a horrible prompt. First of all, using king Philippe I. is against the ToS, so let's go with a generic "king". Let's not confuse the AI with "buts", just say that he is giving the speech to cucumbers. Lastly, specify some style, because this would probably not work out as a photo. My single try is not bad at all and it could definitely be improved. https://labs.openai.com/s/3OUmUxKefJCeLhAk4hkeKX4V
On the other hand, it is more difficult to get it to produce absurd results like these.
my prompt: King Philippe I. of Belgium giving a speech surrounded by [[[[large green vertical cucumbers]]]], digital art in the style of Greg Rutkowski