Live data from Hacker News

Ask HN: DALL-E was trained on watermarked stock images?

news.ycombinator.com

181–190 of 233 posts

Re: Ask HN: DALL-E was trained on watermarked stock images?

#182
post #127

Reminds me of the discussion about GitHub Copilot using the entirety of GitHub as training data. I was honestly baffled how many people, even experts in the field, saw use as training data as non-infringing. With the corrolay that it's apparently perfectly legal to "copyright-wash" a work by feeding it to an AI and have that AI generate a slightly different but extremely similar work. Considering how strict and heavy…

"Copyright washing" seems a lot like clean room reverse engineering to me; this is usually done by having one person read the copyrighted code and describe what it does to another person, who then designs an implementation based on the description.

At least, I can't see a substantial difference in the result.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#183
post #30

I am not a lawyer, but I've had to argue about copyright with several. In the United States, there are two bits of case law that are widely cited and relevant: In Kelly v. Arriba Soft Corp (9th), found that making thumbnails of images for use in a search engine was sufficiently "transformative" that it was ok. Another case, Perfect 10 (9th), found that thumbnails for image search and cached pages were also transforma…

Search engines don't create market harm for a work because they don't compete with it. In fact, they do the opposite: they advertise the work, making it more accessible and increasing exposure. These AI tools on the other hand seem to do the exact opposite. They can (or could, if they got good enough) absolutely compete with a work, and therefore seem like they create substantial market harm. The character of use als…

> Search engines don't create market harm for a work because they don't compete with it. In fact, they do the opposite: they advertise the work, making it more accessible and increasing exposure.

AMP, snippets, Knowledge Base and in-app browsers would like to have a word with you

Re: Ask HN: DALL-E was trained on watermarked stock images?

#184
post #90
post #74

Earlier quoted context omitted.

I'm still not seeing the "transformative" argument: the point of transformation isn't "it is in a different format" but (to quote Wikipedia, which is, of course, dumb... I'm sorry ;P) where one "builds on a copyrighted work in a different manner or for a different purpose from the original". The reason a search engine thumbnail is transformative isn't because it has been transformed to make it smaller... it is becaus…

Autogenerated, often fantastical, never-seen-before AI images strike me as a paradigmatically 'transformative' use. It's novel. It's shocking to many practicioners how flexible & high-quality the images can be. It will unlock all sorts of new downstream creation. The representation that feeds the generation is statistical, even to the point of being plausibly factual : these things/people/places/concepts can be abstr…

I'd argue that if an artefact such as a watermark is copying even more substantially than any other human would and that human would at best be labelled as unoriginal, or doing very derivative work or be in violation of copyright.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#185

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

As a private citizen, what can I do to hasten this trial in court? I have a feeling that the sooner the better, before this becomes a big industry that the US government might not want to hurt.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#186
People here, as always, get hung up on legalese bullshit, but miss the overall picture.

The dynamics in play is highly questionable. Countless artists and photographers put effort into creating their works. They put they work online to get some attention and recognition. A company comes along, scrapes all of it and starts selling access to the model to generate something that looks highly derivative. The original cohort of artists and photographers not only get zero money or attention from this new endeavor, they are now in competition with the resulting model.

In short, someone whose work was essential to building a thing gets no benefits and possibly even gets (financially) harmed by that thing. Just because this gets verbally labeled "fair use" doesn't make it fair.

Additional point:

Just a few years ago a bunch of tech companies were talking about "data dignity". Somehow, magically, this (marketing) term is no longer used anywhere.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#187

Earlier quoted context omitted.

Search engines don't create market harm for a work because they don't compete with it. In fact, they do the opposite: they advertise the work, making it more accessible and increasing exposure. These AI tools on the other hand seem to do the exact opposite. They can (or could, if they got good enough) absolutely compete with a work, and therefore seem like they create substantial market harm. The character of use als…

> Search engines don't create market harm for a work because they don't compete with it. In fact, they do the opposite: they advertise the work, making it more accessible and increasing exposure. AMP, snippets, Knowledge Base and in-app browsers would like to have a word with you

Knowledge Base I grant you, but snippets are a crucial feature to trust a result is correct before clicking through.

AMP is completely unrelated so I'm not sure why you mention it. Website owners have to create a specific version of their own site for AMP to even work.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#188

People here, as always, get hung up on legalese bullshit, but miss the overall picture. The dynamics in play is highly questionable. Countless artists and photographers put effort into creating their works. They put they work online to get some attention and recognition. A company comes along, scrapes all of it and starts selling access to the model to generate something that looks highly derivative. The original coh…

I'm concerned that, and predict that, we will continue to see legal efforts from large data companies to prevent their own data from being used to train similar models. They can use our data, but we can't use theirs. Time will tell.

I also fear our governments are incapable of acting on behalf of the people (non-corporations) in this matter.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#189
post #30

I am not a lawyer, but I've had to argue about copyright with several. In the United States, there are two bits of case law that are widely cited and relevant: In Kelly v. Arriba Soft Corp (9th), found that making thumbnails of images for use in a search engine was sufficiently "transformative" that it was ok. Another case, Perfect 10 (9th), found that thumbnails for image search and cached pages were also transforma…

[deleted]

Re: Ask HN: DALL-E was trained on watermarked stock images?

#190
post #86

Earlier quoted context omitted.

These AI generated images are directly competing with stock images. AI tools are selling images to blogs and other customers that often would purchase stock images instead. The "character of use" is not in favor of dall-e, it is a commercial use. Copyright law does not require getty to block a user agents or ask them not to include their images. Another issue here is that removing copyright management info like a wat…

Whether something is directly competing for the same business would have to be evidenced, and copyright doesn't mean protection from all possible competition - it's just one factor weighed. And fair use protects many commercial uses, too, depending on proportion/character-of-original/etc. But also, none of these images are direct, or even necessarily subtantial, "copies" of other images. The generator learned from ot…

"The generator learned from other images – the same as any human artist might."

A lot of people seem to make this comparison, but I don't think it's fair. It's wrong. A computer is capable of ingesting/processing and "learning" from images at a rate no human can possibly come close to matching. To elaborate, it is not actually learning in the way we normally think of it, as its "brain" is completely different from a human's brain. It is doing something entirely different that should have its own word. Human artists learn from other human artists' work. An AI does something else.

It's also worth noting that the art the AI was trained on was posted online when the technology didn't exist (or if it did in some form it was not in the state it is in now). So an artist having posted their art online for public consumption can't be equated with somehow consenting to its consumption by a web scraper / AI.

Post reply on HN