Live data from Hacker News

Ask HN: DALL-E was trained on watermarked stock images?

news.ycombinator.com

31–40 of 233 posts

Re: Ask HN: DALL-E was trained on watermarked stock images?

#31
All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c....

When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Supreme Courts as it doubtless will), then Copilot is dead, DALL·E is dead, GPT-3 is dead, all of these things will be immediately discontinued in at least the affected jurisdictions, at least until such a time as they get the laws changed or judgements overturned.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#32

> but surely you can't just... use stock photos without paying for the license? They aren't hosting the infringing content. Training on the data is probably covered under fair use. Generations are of _learned_ representations of the dataset, not the dataset itself. This makes it closer to outputting original works (probably owned by the person who used the model). The players involved here are known for being litigio…

I would be very surprised if OpenAI paid anything for these, because it would set precedent that copyright infringement was applicable, which would be fatal down the road. (The only argument they could possibly mount in their defence would be that they wanted to train on the original images without watermarks.)

Re: Ask HN: DALL-E was trained on watermarked stock images?

#33
post #30

I am not a lawyer, but I've had to argue about copyright with several. In the United States, there are two bits of case law that are widely cited and relevant: In Kelly v. Arriba Soft Corp (9th), found that making thumbnails of images for use in a search engine was sufficiently "transformative" that it was ok. Another case, Perfect 10 (9th), found that thumbnails for image search and cached pages were also transforma…

Search engines don't create market harm for a work because they don't compete with it. In fact, they do the opposite: they advertise the work, making it more accessible and increasing exposure.

These AI tools on the other hand seem to do the exact opposite. They can (or could, if they got good enough) absolutely compete with a work, and therefore seem like they create substantial market harm. The character of use also seems vastly different; AI tools are creating images explicitly to be consumed, vs a search engine is basically just an index, and only shows the image in so far as it needs to make it discoverable.

So three of the four tests for fair use seem clearly against AI image generation, at least to me. The only test that possibly goes in favor of AI is the amount or substantiality of copying, but AIs can easily reproduce images, or if not entire images, other substantial subsets of a composition.

I just don't get how these could possibly be fair use.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#34
post #6

I remember when people used to say ianal. Innocent times when we thought there was an objective law and lawyers knew it. But that's not how these things work. The truth is that no one knows. Ultimately a bunch of people will decide how they feel about it. Well-read legal scholars trying really hard to be fair, but still just people. No one can predict with full certainty which way it will go.

>>No one can predict with full certainty which way it will go.

Until somebody tries to float a trial balloon (case) in court.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#35

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

To me this feels like the argument that we should allow Uber and Airbnb because they're sufficiently "transformative" use cases. When clearly they are playing fast and loose by the rules and have taken advantage of being early enough to do so. As soon as the rulemakers caught up, it became obvious that they didn't have a license to operate differently from everyone else, just because they're new and popular.

Personally, I agree that a strike against AI fair use would kill these current generation of tools. But I don't see why that would be the end of it. What it would do is to create a market for open source data sets with liberal licenses. We'd lose something by not being able to train models on every piece of media that has ever been on the internet anywhere, but it's not obvious to me that was ever really reasonable in the first place. If the only way to make AI that can produce good writing is to train it on every piece of writing ever produced in the history of the human race... aren't we missing something? Surely if AI has a future, it'll have to overcome this at some point.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#36
post #26

Earlier quoted context omitted.

yeah except this artist wont go around painting watermarks

Really? https://www.moma.org/learn/moma_learning/andy-warhol-campbel... (Yes I get it's not technically a watermark, but it certainly qualifies as a trade mark in a similar fashion)

my point being if he tries to draw a dog his idea of a dog will not contain the Getty Images watermark and even if he did he would just not draw it

Re: Ask HN: DALL-E was trained on watermarked stock images?

#37

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

My feeling is that while these things may be technically falling under fair use, I really feel like they are running roughshod over a lot of ethical and moral lines and that perhaps "fair use" needs to be redefined to explicitly exclude this kind of processing.

And if it kills these things, oh well. "Being an artist" is a precarious enough existence in this world as is, I'd be delighted to stop worrying about having to compete with an endless sea of algorithmically-generated barely-good-enough spam.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#39

These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.

Prompt Engineering can help a lot but yes, you're basically right: People are generating many, many images and sharing only the best ones with the fewest artifacts.

For simple prompts with little additional guidance, all the diffusion image generators I've seen/used will produce output about like what the author linked most of the time. There are always a few gems, and honing in via prompt engineering helps immensely.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#40

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

Great points but scary. If training ML models on copyrighted data becomes illegal in the US but remains legal in say China or Russia then the US will quickly fall Behind on ML capabilities - major national security implications at the very least. I suspect if the decision went the way you suggest congress would have to change the law to allow training.
Post reply on HN