Live data from Hacker News

Ask HN: DALL-E was trained on watermarked stock images?

news.ycombinator.com

131–140 of 233 posts

Re: Ask HN: DALL-E was trained on watermarked stock images?

#131

Earlier quoted context omitted.

Nah, you could zap the training sets tomorrow and start over with public domain material and it would be fine. In fact I think you could easily get paid to generate more content for it.

For text (the GPT-3 case), that’d work to train a model that had no knowledge of the last century of popular culture or idiom, and was significantly biased to more formal and traditional writing styles. The effects of this would be really quite interesting, but I think it would significantly limit the places it could be usefully applied. For DALL·E and Copilot, I’m confident that you couldn’t find anywhere near enoug…

IANAL, but what you should be able to do is have a set of quotes of one or two sentences each from various sources (books, TV, movies, etc.) that have the modern word or idiom you are specifying. That should then be small enough to fall under fair use (like IMDB, wikiquote, etc. have quotes from films, and good reads, dictionaries, etc. have quotes from books/text), and be plenty enough to capture the meaning of the words/idioms.

You could create your own sentences that you control the copyright of containing the word or idiom, as those words and idioms themselves are not copyrightable. For example: "I fracking hate ice cream!"

For the rest, there is a lot of text upto 1926 (depending on when the author died) that is available for use, so you only need to capture words and idioms changed since then, including any pop culture terms.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#132
post #127

Reminds me of the discussion about GitHub Copilot using the entirety of GitHub as training data. I was honestly baffled how many people, even experts in the field, saw use as training data as non-infringing. With the corrolay that it's apparently perfectly legal to "copyright-wash" a work by feeding it to an AI and have that AI generate a slightly different but extremely similar work. Considering how strict and heavy…

These loopholes are purely theoretical until tested in court. At some point a generating AI will hurt the wrong company, and they will either make a public spectacle out of it in court, or if they see no chance of winning lobby congress to introduce laws that make the case winnable.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#133
post #30

I am not a lawyer, but I've had to argue about copyright with several. In the United States, there are two bits of case law that are widely cited and relevant: In Kelly v. Arriba Soft Corp (9th), found that making thumbnails of images for use in a search engine was sufficiently "transformative" that it was ok. Another case, Perfect 10 (9th), found that thumbnails for image search and cached pages were also transforma…

It seems it is possible to generate images which are very similar to the existing stock photos if you feed getty images' description into DALL-E.

I tried it with a distinctive banana image:

https://imgur.com/a/0OrIr6e

Re: Ask HN: DALL-E was trained on watermarked stock images?

#135

The first thing that I try after generating an image from DALL-E is using reverse image search. I do it on every image that I intend to use, more often than not, I find a very similar image, in this case I discard it and vary my prompts.

> more often than not, I find a very similar image

Can you give an example? I was also doing reverse image searches, and I havent seen a single case of an image being closely related to another unless it was used as the base for inpainting.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#136

Is there a copyright protection in terms of consuming a copyright-protected image? I thought it was only for the purpose of displaying that image. If you're reading the file and reading the data, but not displaying it, is that also protected?

Copyright, as the name implies, is mostly for restricting copying (as in printing copies of a book), but also restricts distribution, adaptation, display, and public performance of that work. In the case of AI, it’s the “adaptation” part which is up for debate. If a person uses an image as part of a training set for an AI image generator, and then uses said AI to generate new images, are those images “adaptations” of the images in the training set? I would suggest that the answer is yes, but current behavior by AI vendors are not in concordance with that view.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#137

Earlier quoted context omitted.

I find Midjourney to be biased towards an artistic representation (for some definition of artistic) When Dall-e is happy to produce children's scribbles or poor imitations.

Try 'poorly drawn ... by a 5 year old using crayons' in Midjourney.

Even then Midjourney is more high-quality :)

See https://imgur.com/gallery/U5zJMcU

Comparison of two prompts, "poorly futuristic landscape by a 5 year-old" and "poorly drawnn highly detailed futuristic landscape dotted by mahcinery and tall buildings by a 5 year-old"

Also, https://imgur.com/gallery/jvEClos

Comparison of "poorly drawn red sports car in the street of a city by a 5 year-old"

Edit: forgot about crayons :D

Re: Ask HN: DALL-E was trained on watermarked stock images?

#138
post #52

These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.

I've been reading some folks saying that "prompt engineering" is a legit future vocation in a world where AI has taken over a lot of creative work And from my experience getting high-quality output from AIs takes a bit of finesse. Not quite unlike crafting a good Google query so... yes

Is command language the ultimate interface? I doubt it. Similar to how GUI supersedes CLI in most use cases, we should be able to indicate "warmer" / "colder" preference to generate new images from previous attempts.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#139

All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…

I have been saying for a long time. These big companies with an investment in getting these systems working could collaborate with twitter and instagram to include image license options. They could give users the option to opt-in to data collection. This would help the public feel like they’re giving consent, and it would also lead to the creation of massive, properly licensed open data sets. There’s already some big…

> They could give users the option to opt-in to data collection. This would help the public feel like they’re giving consent

Would we get paid? I'd care a lot less about OpenAI profiting from my work if I was getting commission every time my work was hit in the training data.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#140
I'm finding it amusing that everyone immediately assumes infringement, OpenAI is a company that will not be inviting lawsuits.

We can't assume any licensing behind closed doors, my guess is that OpenAI has an agreement with Getty, take a look at the licensing in this Observer piece, it's been licensed by Getty, this would indicate that Getty are happy with scraping.

https://www.theguardian.com/commentisfree/2022/aug/20/ai-art...

Besides, this is not infringement in principle, the AI has been trained to think that high-quality news images have watermarks.

Post reply on HN