Live data from Hacker News

Ask HN: DALL-E was trained on watermarked stock images?

news.ycombinator.com

71–80 of 233 posts

Re: Ask HN: DALL-E was trained on watermarked stock images?

#71

Earlier quoted context omitted.

Top 1% is a bit exaggerated, but there is definitely a lot of not good stuff. I find that Dall-E does especially poorly with underspecified prompts too, unlike something like Midjourney which can give visually pleasing photos for even the most abstract concepts. Dall-E tends to do better with concrete and specific prompts. Here's an example: Stressful Shapes Dall-E: https://i.imgur.com/JBkSh0y.png Midjourney: https:/…

I gather Midjourney was trained primarily using Journey album covers?

https://laion.ai/blog/laion-aesthetics/

Re: Ask HN: DALL-E was trained on watermarked stock images?

#72
post #52

These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.

I've been reading some folks saying that "prompt engineering" is a legit future vocation in a world where AI has taken over a lot of creative work And from my experience getting high-quality output from AIs takes a bit of finesse. Not quite unlike crafting a good Google query so... yes

doubtful - I'm tuning GPT3 with good midjourney prompts as we speak

Re: Ask HN: DALL-E was trained on watermarked stock images?

#73

Earlier quoted context omitted.

Nah, you could zap the training sets tomorrow and start over with public domain material and it would be fine. In fact I think you could easily get paid to generate more content for it.

For text (the GPT-3 case), that’d work to train a model that had no knowledge of the last century of popular culture or idiom, and was significantly biased to more formal and traditional writing styles. The effects of this would be really quite interesting, but I think it would significantly limit the places it could be usefully applied. For DALL·E and Copilot, I’m confident that you couldn’t find anywhere near enoug…

There is tons of english text and images with permissive licensing. All of stack overflow, wikipedia is creative commons. Anything created by the US government or many other governments is public domain

Re: Ask HN: DALL-E was trained on watermarked stock images?

#74
post #48

Earlier quoted context omitted.

Search engines don't create market harm for a work because they don't compete with it. In fact, they do the opposite: they advertise the work, making it more accessible and increasing exposure. These AI tools on the other hand seem to do the exact opposite. They can (or could, if they got good enough) absolutely compete with a work, and therefore seem like they create substantial market harm. The character of use als…

As I see it, 3 of the 4 tests are strongly in OpenAI's favor; the 'market effect' is mixed. (1) The use is highly transformative; (2) the images used were offered to the anonymous browsing public (with watermarks); (3) the end effect of training will only retain a tiny spectral distilled essence of any individual photo, or even a giant source corpus; (4) there's a potential risk of market competition from the ultimat…

I'm still not seeing the "transformative" argument: the point of transformation isn't "it is in a different format" but (to quote Wikipedia, which is, of course, dumb... I'm sorry ;P) where one "builds on a copyrighted work in a different manner or for a different purpose from the original". The reason a search engine thumbnail is transformative isn't because it has been transformed to make it smaller... it is because the purpose of the resulting use of the image is somewhat unrelated to the use the original author was going for when they made the original image. At issue here is then that, rather than using an original image from Getty Images, someone decided to take all of the images from Getty Images and churn them through some algorithm that generated an image that directly competed with the original images from Getty Images. So like, sure: if you really only narrowly want to talk about OpenAI, what they are themselves doing (training and distributing a model) might potentially be legal, but the people using the result would seem to be in serious hot water... oh, and actually, I think they run it all a service, don't they? So no: I don't even think that defense works, as OpenAI is in some sense not even selling a model, they are merely directly competing with Getty Images to provide sell photos to people.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#75

What is interesting is a human analogy. Say you were an artist who went to every art show and museum and studied all the art there. If you produced a work of art solely from memory that contained large portions of other people's copyrighted art, would that still fall under copyright/require licensing?

Alternatively, if you memorized some GPL code, can you write a copy of it and put in a proprietary licence?

Re: Ask HN: DALL-E was trained on watermarked stock images?

#76

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

I personally feel like that lawsuit happens the moment someone builds the version of this that works on music; in my experience arguing before the copyright office at the library of congress, the people who tend to be the most omnipresent is the RIAA, and when someone releases an AI-generated piece of music that sort of sounds like some recent Taylor Swift song but using that infamous sample from Under Pressure / Ice Ice Baby, the lawsuit will be filed within days.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#77

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

To me this feels like the argument that we should allow Uber and Airbnb because they're sufficiently "transformative" use cases. When clearly they are playing fast and loose by the rules and have taken advantage of being early enough to do so. As soon as the rulemakers caught up, it became obvious that they didn't have a license to operate differently from everyone else, just because they're new and popular. Personal…

> As soon as the rulemakers caught up, it became obvious that they didn't have a license to operate differently from everyone else, just because they're new and popular.

This is well said. One of the primary advantages of these businesses was evading the regulation and taxation that their competitors were subject to.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#78

Earlier quoted context omitted.

To me this feels like the argument that we should allow Uber and Airbnb because they're sufficiently "transformative" use cases. When clearly they are playing fast and loose by the rules and have taken advantage of being early enough to do so. As soon as the rulemakers caught up, it became obvious that they didn't have a license to operate differently from everyone else, just because they're new and popular. Personal…

> As soon as the rulemakers caught up, it became obvious that they didn't have a license to operate differently from everyone else, just because they're new and popular. This is well said. One of the primary advantages of these businesses was evading the regulation and taxation that their competitors were subject to.

Indeed, but this is the proof the market didn't want the regulation and taxation and wanted better technology.

Shouldn't we listen to what people want over bureaucrats?

Re: Ask HN: DALL-E was trained on watermarked stock images?

#79

Earlier quoted context omitted.

For text (the GPT-3 case), that’d work to train a model that had no knowledge of the last century of popular culture or idiom, and was significantly biased to more formal and traditional writing styles. The effects of this would be really quite interesting, but I think it would significantly limit the places it could be usefully applied. For DALL·E and Copilot, I’m confident that you couldn’t find anywhere near enoug…

There is tons of english text and images with permissive licensing. All of stack overflow, wikipedia is creative commons. Anything created by the US government or many other governments is public domain

The terms of the CC-BY-SA licenses that Stack Overflow and Wikipedia largely use cannot practically be satisfied in a data model. By design, all outputs derive from all sources to some extent, and the licensing requires that they generally be specifically identified, so you can’t just say “from Wikipedia” or “from Stack Overflow” but “from such-and-such a page, by so-and-so”.

“Permissive” is not enough. You need no-strings-attached, and attribution is a string. Hence mostly talking about public domain materials, which make up the vast majority of suitable materials.

I suspect that covered works of the USA federal government would be quite a large fraction of the public domain material (as reckoned by the USA) from the last 70 years. I don’t believe it’d be enough to be particularly useful, certainly not for pop culture knowledge or colloquial idiom.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#80

Earlier quoted context omitted.

I have been saying for a long time. These big companies with an investment in getting these systems working could collaborate with twitter and instagram to include image license options. They could give users the option to opt-in to data collection. This would help the public feel like they’re giving consent, and it would also lead to the creation of massive, properly licensed open data sets. There’s already some big…

I’m not particularly familiar with open image data sets, but I doubt that most of them are suitable. Open Images , for example, uses CC-BY images ( https://storage.googleapis.com/openimages/web/factsfigures.h... ). Without the fair use exemption, this would suggest that if you used a model trained on that, you would have to comply with the license of every image, which would mean providing attribution for every singl…

Actually, I think they could comply by providing a list of every author in the dataset, though this is following the letter rather than the spirit.

They could also do research on ways to get the model to return the top ten influential works for some output, and make a legal argument that this is a best effort given technical challenges with tracing every source.

For the first point, I get that by reading the license at this image:

https://commons.m.wikimedia.org/wiki/File:3H8A7368.jpg

“attribution – You must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.

share alike – If you remix, transform, or build upon the material, you must distribute your contributions under the same or compatible license as the original.”

So, output has to have the same license and worst case the image is accompanied by a link to a list of every author in the dataset. This is a far cry from the “this research would be useless if copyright was enforced” as some people suggest.

Here’s an open dataset I found which does not require attribution: https://www.pexels.com/creative-commons-images/

And Wikimedia commons, some of which require attribution: https://commons.m.wikimedia.org/wiki/Category:Images

And it is easy to go take a 4k video camera and start collecting tens of thousands of frames of your own images.

My point is that people are throwing up their hands saying oh, respecting the copyright of artists is impossible. But it feels very unfair that these huge companies are walking all over the copyright of small artists, but if we took their code to re use their lawyers would sink us overnight. This upsets a lot of people and it’s a bad look. I don’t actually like copyright but if everyone else has to follow the rules I don’t like giving them a free pass.

Post reply on HN