Live data from Hacker News

Ask HN: DALL-E was trained on watermarked stock images?

news.ycombinator.com

51–60 of 233 posts

Re: Ask HN: DALL-E was trained on watermarked stock images?

#51
post #43
post #41

Earlier quoted context omitted.

Of course people are more likely to share the best iamges – or in this case, the one most illustrative of their concern (about watermarks). Also: my sense is that getting the best results often requires a lot of extra coaching with style/detail words. As we can't see the prompt here, we don't know what sort of style/details were requested. GIGO.

You're right. This shows the prompt and it doesn't have such style directives https://ibb.co/gz5RDkB

Also, a construction like 'but' that tries to override another expectation may be suboptimal. I gave the same concept a few tries, with more 'sweeteners'. First batch, for prompt "news photo of the King of Belgium giving a speech to an audience that is entirely cucumbers, award-winning, well-composed, detailed surroundings" – & it's a bit better:

https://labs.openai.com/s/9YF5WxF1GoZVdzpLBAQYp2Zg

https://labs.openai.com/s/wBfHevs9hIZXvzkJ686mFmn3

https://labs.openai.com/s/M0i029fZnYjQXHFobUpw7eun (best of batch imo)

https://labs.openai.com/s/Hf4z0M9M3KBr6IaaKsjEt9Mx

A few more tries didn't manage to create any photorealistic shots with actual cucumbers-in-seats – perhaps due to the absurd contrasts required – but shifting to a 'cartoon' style with the prompt "editorial cartoon of the King of Belgium giving a speech to many cheering cucumbers, professional illustrator" got a lot closer:

https://labs.openai.com/s/4IonSKYkl0okhNvzJmEAH30K (good)

https://labs.openai.com/s/ZeadCzZ9WqeASYXPlOb13wDV (good)

https://labs.openai.com/s/nERf6bALKEBsQQBvPVsAH7o4 (good)

https://labs.openai.com/s/7dsZu3bwtfGZZxxJYtg9lTVf

If I had more time & credits to burn, I suspect working off those could eventually hit something really apt... but it takes some work & tinkering.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#52

These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.

I've been reading some folks saying that "prompt engineering" is a legit future vocation in a world where AI has taken over a lot of creative work

And from my experience getting high-quality output from AIs takes a bit of finesse. Not quite unlike crafting a good Google query

so... yes

Re: Ask HN: DALL-E was trained on watermarked stock images?

#54

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

It's not going to fail: the US courts are big company biased, and all the big companies are going to show out in force and money to ensure they get the result they want.

But even extending that: knocking copyright'd images out isn't going to stop these systems. We know they work now, so if you have to be careful about licensing then that's just going to be done.

The idea that any of these platforms will "die" if copyright fair use doesn't automatically apply is magical thinking. Most art is worthless - companies hoovering up huge corpuses with the correct rights assignment for machine learning is going to be the new business.

A company like Disney will drop every piece of output from their staff into a dataset for Disney, then license it under terms to other companies - the tech works, so "invent me a disney character looking like..." would have value internally, just to Disney, for idea generation and refinement - arguably a lot more then to anyone else because they would still retain the artist resources to capitalize on it.

Right now, a bunch of people who told themselves that despite the pay, they weren't going to be replaced by AI are shrieking that it's turned out not to be the case (it was obvious for a few years something like this was coming though). They're reaching for every legal tool that they hope will kill these things, forgetting that it's never worked out like that. Copyright being a problem when it happens to you as an individual, is different to when it happens to MegaCorp Inc. which is constantly being sued, has limited liability, and puts payouts down as a line-item expense.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#55
Regardless of whether or not training an AI on stock images violates the license, there's a very real problem with that watermark being present, which is that it proves their AI is prone to copying large swaths of images from gettyimages unaltered, and that definitely is a license violation.

This makes me think back to the controversy over github copilot; if these AIs are going to be trained on other peoples' IP then somebody needs to be held accountable when they commit plagiarism.

Otherwise, im sure Microsoft won't mind my new "gamemaker AI" that i trained on that new halo game last year, or this "OS AI" that I trained on windows 11.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#56

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

Nah, you could zap the training sets tomorrow and start over with public domain material and it would be fine. In fact I think you could easily get paid to generate more content for it.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#57

Kids in school are also trained on stock images https://www.reddit.com/r/KidsAreFuckingStupid/comments/8tgxs...

That's technically not a stock image, it's a portait that has been public domain for a long time.

But you've seen many PD images reshared by stock imagery companies. It raises the question of why false assertions of ownership aren't easily prosecuted, given that they constitute a kind of fraud upon the public.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#58

> but surely you can't just... use stock photos without paying for the license? They aren't hosting the infringing content. Training on the data is probably covered under fair use. Generations are of _learned_ representations of the dataset, not the dataset itself. This makes it closer to outputting original works (probably owned by the person who used the model). The players involved here are known for being litigio…

What if I write a machine learning algorithm that only generates images that it has seen in the training dataset, with one pixel slightly different.

It won't be transformative enough and you'd probably lose the case.

(IANAL)

Re: Ask HN: DALL-E was trained on watermarked stock images?

#59
post #46
post #40

Earlier quoted context omitted.

Great points but scary. If training ML models on copyrighted data becomes illegal in the US but remains legal in say China or Russia then the US will quickly fall Behind on ML capabilities - major national security implications at the very least. I suspect if the decision went the way you suggest congress would have to change the law to allow training.

Isn't that true for all technology? In the U.S. we have the specwriter system which leads to inefficiency to get around copyright. In China or Russia they just copy the code and iterate.

As a Russian programmer, worked in companies big and small, I cannot say this is even remotely right.

All companies I worked with really cared about cleanness of the origin of the code.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#60
post #41

These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.

Of course people are more likely to share the best iamges – or in this case, the one most illustrative of their concern (about watermarks). Also: my sense is that getting the best results often requires a lot of extra coaching with style/detail words. As we can't see the prompt here, we don't know what sort of style/details were requested. GIGO.

OP did say what prompt they used
Post reply on HN