Live data from Hacker News

Ask HN: DALL-E was trained on watermarked stock images?

news.ycombinator.com

171–180 of 233 posts

Re: Ask HN: DALL-E was trained on watermarked stock images?

#171
post #148
post #86

Earlier quoted context omitted.

Whether something is directly competing for the same business would have to be evidenced, and copyright doesn't mean protection from all possible competition - it's just one factor weighed. And fair use protects many commercial uses, too, depending on proportion/character-of-original/etc. But also, none of these images are direct, or even necessarily subtantial, "copies" of other images. The generator learned from ot…

It’s going to be interesting what the stock companies will do. Maybe they will make their own Image Generator. Perhaps we will see a case based on the new factor that is AI. An AI is not artist; they can’t be conflated. A decent artists can churn out maybe 5-10 works if he is productive. AI can churn out by the hundreds or thousands if needed. The process also isn’t the same. Anyway it will be interesting to watch th…

AI generated images cant be copyrighted.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#172

> but surely you can't just... use stock photos without paying for the license? They aren't hosting the infringing content. Training on the data is probably covered under fair use. Generations are of _learned_ representations of the dataset, not the dataset itself. This makes it closer to outputting original works (probably owned by the person who used the model). The players involved here are known for being litigio…

if they paid for access, or permission, why train on the watermark versions? I’m guessing they assumed fair use and there will be lawsuits.

Is that representation of the watermark a trademark? If so, then copyright infringement might not matter, but use of the trademark may.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#173

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

They will be discontinued, but of course the profits made during all this time-- with everyone including those companies knowing how it is basically laundering intellectual property-- will stick.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#174
post #30

I am not a lawyer, but I've had to argue about copyright with several. In the United States, there are two bits of case law that are widely cited and relevant: In Kelly v. Arriba Soft Corp (9th), found that making thumbnails of images for use in a search engine was sufficiently "transformative" that it was ok. Another case, Perfect 10 (9th), found that thumbnails for image search and cached pages were also transforma…

From what I understand, the actual process of fair use boils down to "the judge decides in his/her gut if the use is fair, and then writes up the analysis to justify coming to that conclusion." If you look at the recent SCOTUS opinion in Google v Oracle, you can see how two judges can look at the same facts and come to almost diametrically opposed fair use analyses. My further understanding is that generally the #1 overriding concern in fair use analysis is money, which means you're more likely to see analysis along Thomas's dissent than Breyer's opinion.

In this case, let me give a fair use analysis that is going to suggest that this isn't fair. Factor 1 weighs against fair use: it's not transformative because, well, transformative is extremely narrowly interpreted against fair use. Factor 2 weighs against fair use because, well, it's factor 2 and it weighs against fair use unless the underlying copyright was paper-thin in the first place. In factor 3, it's weighing against fair use because it's not copying the minimal amount of the original work to get what it needs (it copied the watermark after all!). And factor 4 of course weighs against fair use because you're essentially creating stock images which is naturally in the exact same market that a stock image provider is in.

If you wanted to write a fair use analysis that finds fair use, you'd argue instead that the work was transformative, and the amount copied also weighs in favor of fair use (thus converting factors 1 and 3 to weigh in favor of fair use). You might try to argue that it's a completely different market, but I'm incredibly skeptical that such an argument could win over both a district court and an appeals court (although Breyer's opinion in Google v Oracle did basically follow this thread of analysis, its repetition is unlikely since everyone wants to pretend that Google v Oracle has 0 impact to anything outside of software). Such an analysis is possible, but unlikely, since the unspoken factor of "could you have paid for this" tends to be the factor that wins out over everything else.

Note that we are going to have a SCOTUS case in the fall that will specifically explore transformative uses in the context of fair use: Warhol v Goldsmith (https://www.scotusblog.com/case-files/cases/andy-warhol-foun...). I'm not going to hold my breath that the use will be found fair, though.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#175

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

To me this feels like the argument that we should allow Uber and Airbnb because they're sufficiently "transformative" use cases. When clearly they are playing fast and loose by the rules and have taken advantage of being early enough to do so. As soon as the rulemakers caught up, it became obvious that they didn't have a license to operate differently from everyone else, just because they're new and popular. Personal…

There is no "fair use" when it comes to laws and regulations.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#176
I don't care much for what laws say. If the only way someones service can work is by ingesting the work of someone else, without compensation, and then compete with that same person, that is wrong.

If a company reverse engineers a competitors product, they still buy the product to tear it apart and figure out how it works.

If a student learns from their teacher, then goes on to sell a similar kind of work as what their teacher makes, at least the student paid for the classes.

This arrangement offers none of that. As long as theft is illegal, this should be. I'd call it parasitic, but it isn't; this is a parasite who's sole intent is to kill the host.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#177
post #106

Earlier quoted context omitted.

Top 1% is a bit exaggerated, but there is definitely a lot of not good stuff. I find that Dall-E does especially poorly with underspecified prompts too, unlike something like Midjourney which can give visually pleasing photos for even the most abstract concepts. Dall-E tends to do better with concrete and specific prompts. Here's an example: Stressful Shapes Dall-E: https://i.imgur.com/JBkSh0y.png Midjourney: https:/…

Now that I'm aware and biased, DALL-E's first image indeed looks very much like stock photo training. This would also make sense given how they can correlate the image with words completely for free due to pretty extensive metadata. What puzzles me is if the Getty Images logo can sometimes appear. If you only have a Getty account, you get rid of the logo and can legally use them royalty free?

No, but you can input your getty image to StableDiffussion img2img and see what's out

Re: Ask HN: DALL-E was trained on watermarked stock images?

#180
post #127

Reminds me of the discussion about GitHub Copilot using the entirety of GitHub as training data. I was honestly baffled how many people, even experts in the field, saw use as training data as non-infringing. With the corrolay that it's apparently perfectly legal to "copyright-wash" a work by feeding it to an AI and have that AI generate a slightly different but extremely similar work. Considering how strict and heavy…

These loopholes are purely theoretical until tested in court. At some point a generating AI will hurt the wrong company, and they will either make a public spectacle out of it in court, or if they see no chance of winning lobby congress to introduce laws that make the case winnable.

Yeah, things should get interesting when the first model makes use of Rings of Power or House of the Dragon footage or whatever the latest superhero movie is.

I wonder if we'll see a "Hollywood vs Silicon Valley" lobbying battle. Or possibly "Amazon media division vs Amazon AI division"...

Post reply on HN