Live data from Hacker News

Ask HN: DALL-E was trained on watermarked stock images?

news.ycombinator.com

201–210 of 233 posts

Re: Ask HN: DALL-E was trained on watermarked stock images?

#201
post #180

Earlier quoted context omitted.

These loopholes are purely theoretical until tested in court. At some point a generating AI will hurt the wrong company, and they will either make a public spectacle out of it in court, or if they see no chance of winning lobby congress to introduce laws that make the case winnable.

Yeah, things should get interesting when the first model makes use of Rings of Power or House of the Dragon footage or whatever the latest superhero movie is. I wonder if we'll see a "Hollywood vs Silicon Valley" lobbying battle. Or possibly "Amazon media division vs Amazon AI division"...

I think silicon valley would win. I saw some analysis a long time a go that basically indicated a couple big companies could likely buy out the entire Hollywood and music industry and fully own them and make most copyright issues go away.

I dont know if its still true but its really a big difference in how much capital, revenue, and money there is. Hollywood is pennies in comparison. I think Big tech would easily win.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#202
post #18

> but surely you can't just... use stock photos without paying for the license? They aren't hosting the infringing content. Training on the data is probably covered under fair use. Generations are of _learned_ representations of the dataset, not the dataset itself. This makes it closer to outputting original works (probably owned by the person who used the model). The players involved here are known for being litigio…

If they had been paying for the images upfront, wouldn't you expect them to train the model on the non-watermarked versions?

The watermarked version might be more prolific with better metatags and descriptions around them.

The non watermarked versions are likely internal only and have far less diverse descriptions.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#204
post #90

Earlier quoted context omitted.

Autogenerated, often fantastical, never-seen-before AI images strike me as a paradigmatically 'transformative' use. It's novel. It's shocking to many practicioners how flexible & high-quality the images can be. It will unlock all sorts of new downstream creation. The representation that feeds the generation is statistical, even to the point of being plausibly factual : these things/people/places/concepts can be abstr…

"AI artist" doesn't add any of its "own flair". It builds exclusively on past experience and work of humans. And it also directly completes with them without any thought of credit or compensation. People are really underplaying how damaging this is going to be for the industry. It's going to completely decimate it. You can already see people using names of artists in the DALL-E prompt to get "their" work for few doll…

One could make the same case about humans, nobody works in a vacuum. Even though he used it in a pejorative sense, Sir Isaac Newton, the famous English scientist, once said, “If I have seen further, it is by standing on the shoulders of giants.”

That humans are capable of developing their own style could still be argued that it's just a intermixing of previous work that they've seen, but they've combined it in a different way, which effectively is exactly what these generative systems do.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#205
post #110
post #99

Earlier quoted context omitted.

Op constructed a horrible prompt. First of all, using king Philippe I. is against the ToS, so let's go with a generic "king". Let's not confuse the AI with "buts", just say that he is giving the speech to cucumbers. Lastly, specify some style, because this would probably not work out as a photo. My single try is not bad at all and it could definitely be improved. https://labs.openai.com/s/3OUmUxKefJCeLhAk4hkeKX4V

I tried it with Stable Diffusion as well. You can use actual people and the model is even pretty decent at many of the famous ones. On the other hand, it is more difficult to get it to produce absurd results like these. my prompt: King Philippe I. of Belgium giving a speech surrounded by [[[[large green vertical cucumbers]]]], digital art in the style of Greg Rutkowski https://files.catbox.moe/1ej1a4.png

Is the usage of the surrounding brackets some kind of keyword weights specific to stable diffusion?

Re: Ask HN: DALL-E was trained on watermarked stock images?

#206

Earlier quoted context omitted.

You will still be competing against it. It'll just be outsourced and presented as manual work.

True. But there wouldn’t be a growing number of easily accessible websites happily letting everyone type whatever prompt they desire, and there wouldn’t be further development being done in this area. “Zero people trying to disrupt my job with art generators” would be ideal, but “a lot fewer people trying to disrupt my job with illegal art generators they can get in trouble for using” is still better than “hey we rel…

There will be. They'll just be outside of your country's jurisdiction.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#207
post #121

Earlier quoted context omitted.

Top 1% is a bit exaggerated, but there is definitely a lot of not good stuff. I find that Dall-E does especially poorly with underspecified prompts too, unlike something like Midjourney which can give visually pleasing photos for even the most abstract concepts. Dall-E tends to do better with concrete and specific prompts. Here's an example: Stressful Shapes Dall-E: https://i.imgur.com/JBkSh0y.png Midjourney: https:/…

But still "king of belgium giving a speech to an audience, but the audience members are cucumbers" is very specific. And I don't see the king of Belgium anywhere, two pictures have absolutely nothing to do with the prompt (no king, no speech, no audience, no cucumber), one has the speech and audience but no king or cucumber. Graphically, they are deep into the uncanny valley. Only the third image is kind of right, if…

Could that have to do with Sophie Wilmès, Belgian PM 2019–2020? (https://en.wikipedia.org/wiki/Prime_Minister_of_Belgium#Livi...)

Re: Ask HN: DALL-E was trained on watermarked stock images?

#208

Earlier quoted context omitted.

"AI artist" doesn't add any of its "own flair". It builds exclusively on past experience and work of humans. And it also directly completes with them without any thought of credit or compensation. People are really underplaying how damaging this is going to be for the industry. It's going to completely decimate it. You can already see people using names of artists in the DALL-E prompt to get "their" work for few doll…

One could make the same case about humans, nobody works in a vacuum. Even though he used it in a pejorative sense, Sir Isaac Newton, the famous English scientist, once said, “If I have seen further, it is by standing on the shoulders of giants.” That humans are capable of developing their own style could still be argued that it's just a intermixing of previous work that they've seen, but they've combined it in a diff…

Of course humans build on work of other people. And what they do is partialy a mashup. But their work is not only replication of visual patterns. Its thinking its other non visual experiences its their politics and world views combined in their work. Often its their life project.

To think that artists only mash up what was before them is quite obviously wrong.

But its exactly only thing the tech does.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#209
post #90

Earlier quoted context omitted.

Autogenerated, often fantastical, never-seen-before AI images strike me as a paradigmatically 'transformative' use. It's novel. It's shocking to many practicioners how flexible & high-quality the images can be. It will unlock all sorts of new downstream creation. The representation that feeds the generation is statistical, even to the point of being plausibly factual : these things/people/places/concepts can be abstr…

Perhaps I’m misunderstanding your argument, but my counterexample would be: if a human digital artist transformed a Getty image, resulting a fantastical, never-before-seen result, using software like Photoshop, that use would be no more defensible. If anything, the vast scale at which this occurs in AI makes it worse.

I think your hypothetical would depend on the character & extent of the transformation. Mere filters that leave the original recognizable? Probably an infringement. But creative application of transformations to express new ideas? Maybe not – especially if the derivative is a comment/parody on the original, that actually increases interest in it. Most art is a conversation with the past, reusing recognizable motifs & often even exact elements.

For example:

Andy Warhol died in 1987, 35 years ago. One of his 'Prince' collages dating to the early 80s used another photographer's photo, without permission. In 2019, one federal judge ruled that was not infringement. An appeals judge then said it was.

The Supreme Court has decided to take the case.

The US Copyright Office & Department of Justice agree with the photographer in briefs filed with the court... but the mere fact the Supreme Court took the case indicates they think there might be issues with the appeals court ruling. They might agree with the original judge!

Oral arguments come this October. See:

https://www.reuters.com/legal/litigation/us-backs-photograph...

So, when all the (possible) disputes over AI-training-on-copyrighted-images resolve – maybe in the 2030s or 2040s? – what will the laws say, & courts decide? It'll depend a lot on other specifics, & reasoning, that may not be evident now.

Re: Ask HN: DALL-E was trained on watermarked stock images?

#210
post #196

Earlier quoted context omitted.

"AI artist" doesn't add any of its "own flair". It builds exclusively on past experience and work of humans. And it also directly completes with them without any thought of credit or compensation. People are really underplaying how damaging this is going to be for the industry. It's going to completely decimate it. You can already see people using names of artists in the DALL-E prompt to get "their" work for few doll…

The synergistic effect of all the AI's inputs absolutely results in a unique new 'flair', with extensions, reversals, and mash-ups of styles just as in human-made artistic styles. And AI "builds exclusively on past experience and work of humans" just like any young new human artist equally does. In many cases, you can even tell the different models' outputs apart, not by raw quality or glitches, but by hard-to-descri…

Indeed the genie is out. And while we will get some interesting AI uses ultimately this is degenerative tech. In the end we end up with less authentic, less unpredictible and less delightfull art. Instead we get the perfectly suited to us, predictible, mediocre stuff.

I said it in comment above - yes people build on work of others but they also bring lots of their originality and intelect. Part of what people do is truly uniquely theirs and piece by piece we progress as a whole.

The crutial detail is that AI learns only from visual patterns from past and cant think at all. And humans learn from everything around them and think about it deeply.

Post reply on HN