Live data from Hacker News

Ask HN: DALL-E was trained on watermarked stock images?

news.ycombinator.com

61–70 of 233 posts

Re: Ask HN: DALL-E was trained on watermarked stock images?

#61

Earlier quoted context omitted.

What if I write a machine learning algorithm that only generates images that it has seen in the training dataset, with one pixel slightly different.

It won't be transformative enough and you'd probably lose the case. (IANAL)

What about two pixels?

Re: Ask HN: DALL-E was trained on watermarked stock images?

#62

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

Nah, you could zap the training sets tomorrow and start over with public domain material and it would be fine. In fact I think you could easily get paid to generate more content for it.

For text (the GPT-3 case), that’d work to train a model that had no knowledge of the last century of popular culture or idiom, and was significantly biased to more formal and traditional writing styles. The effects of this would be really quite interesting, but I think it would significantly limit the places it could be usefully applied.

For DALL·E and Copilot, I’m confident that you couldn’t find anywhere near enough material to produce results anywhere near as good as what there is now. I strongly suspect the results would be too poor to be useful in most places where they may be useful now.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#63

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

I have been saying for a long time. These big companies with an investment in getting these systems working could collaborate with twitter and instagram to include image license options. They could give users the option to opt-in to data collection. This would help the public feel like they’re giving consent, and it would also lead to the creation of massive, properly licensed open data sets. There’s already some big open image datasets to start with.

And of course, they could always work more on sample efficiency.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#64
post #54

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

It's not going to fail: the US courts are big company biased, and all the big companies are going to show out in force and money to ensure they get the result they want. But even extending that: knocking copyright'd images out isn't going to stop these systems. We know they work now, so if you have to be careful about licensing then that's just going to be done. The idea that any of these platforms will "die" if copy…

For my part, I think you’re right, and that it’s unlikely to fail in at least the USA, and, after consideration of what you wrote, that if it did fail in any meaningful way it would push things in the direction of copyright pools (like patent pools); but at the very least, it would be a massive disruption which would take some time to be sorted out and require a certain degree of starting from scratch in data sets; and all up it’d probably favour big business even more heavily than the current informal consensus. I think there’s also a fair chance in such a situation that European countries with their different approach to copyright philosophy would act as a balancing force, striking down overly-general copyright-assignment-equivalent clauses in terms of service and the likes, which would be the only real way of sucking up as much everything as these models need to work well, especially for providing retroactive relicensing (to avoid a big hole in their sources).

Re: Ask HN: DALL-E was trained on watermarked stock images?

#65

What is interesting is a human analogy. Say you were an artist who went to every art show and museum and studied all the art there. If you produced a work of art solely from memory that contained large portions of other people's copyrighted art, would that still fall under copyright/require licensing?

yeah except this artist wont go around painting watermarks

Copying signatures one entera into forgery territory…

Re: Ask HN: DALL-E was trained on watermarked stock images?

#66

These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.

Top 1% is a bit exaggerated, but there is definitely a lot of not good stuff. I find that Dall-E does especially poorly with underspecified prompts too, unlike something like Midjourney which can give visually pleasing photos for even the most abstract concepts. Dall-E tends to do better with concrete and specific prompts. Here's an example: Stressful Shapes Dall-E: https://i.imgur.com/JBkSh0y.png Midjourney: https:/…

I gather Midjourney was trained primarily using Journey album covers?

Re: Ask HN: DALL-E was trained on watermarked stock images?

#67
post #8

Legally wouldn't it just boil down to the license on the watermarked image? BTW you can add 'royalty free' to the prompt to get rid of those most of the time (lol?).

> royalty free

Wouldn’t that remove the king of Belgium? Or add a “down with the king” placard?

Re: Ask HN: DALL-E was trained on watermarked stock images?

#68

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

I have been saying for a long time. These big companies with an investment in getting these systems working could collaborate with twitter and instagram to include image license options. They could give users the option to opt-in to data collection. This would help the public feel like they’re giving consent, and it would also lead to the creation of massive, properly licensed open data sets. There’s already some big…

I’m not particularly familiar with open image data sets, but I doubt that most of them are suitable. Open Images, for example, uses CC-BY images (https://storage.googleapis.com/openimages/web/factsfigures.h...). Without the fair use exemption, this would suggest that if you used a model trained on that, you would have to comply with the license of every image, which would mean providing attribution for every single item in the set, which is somewhere between infeasible and impractical.

The only types of licenses suitable are ones that require nothing like attribution. This is why you’d mostly be limited to public domain materials (though if it went down this way, you’d find terms of service popping up that included a license grant for model training and selling either your data for model training or trained models without any sort of attribution or remuneration).

Re: Ask HN: DALL-E was trained on watermarked stock images?

#69
post #48

Earlier quoted context omitted.

Search engines don't create market harm for a work because they don't compete with it. In fact, they do the opposite: they advertise the work, making it more accessible and increasing exposure. These AI tools on the other hand seem to do the exact opposite. They can (or could, if they got good enough) absolutely compete with a work, and therefore seem like they create substantial market harm. The character of use als…

As I see it, 3 of the 4 tests are strongly in OpenAI's favor; the 'market effect' is mixed. (1) The use is highly transformative; (2) the images used were offered to the anonymous browsing public (with watermarks); (3) the end effect of training will only retain a tiny spectral distilled essence of any individual photo, or even a giant source corpus; (4) there's a potential risk of market competition from the ultimat…

These AI generated images are directly competing with stock images. AI tools are selling images to blogs and other customers that often would purchase stock images instead.

The "character of use" is not in favor of dall-e, it is a commercial use.

Copyright law does not require getty to block a user agents or ask them not to include their images.

Another issue here is that removing copyright management info like a watermark is a violation of the DMCA, separate from fair use or copyright infringement. These cases have statutory damages and attorneys fees awarded.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#70

These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.

Top 1% is a bit exaggerated, but there is definitely a lot of not good stuff. I find that Dall-E does especially poorly with underspecified prompts too, unlike something like Midjourney which can give visually pleasing photos for even the most abstract concepts. Dall-E tends to do better with concrete and specific prompts. Here's an example: Stressful Shapes Dall-E: https://i.imgur.com/JBkSh0y.png Midjourney: https:/…

I find Midjourney to be biased towards an artistic representation (for some definition of artistic) When Dall-e is happy to produce children's scribbles or poor imitations.
Post reply on HN