Earlier quoted context omitted.
It won't be transformative enough and you'd probably lose the case. (IANAL)
What about two pixels?
Ask HN: DALL-E was trained on watermarked stock images?
211–220 of 233 posts
Re: Ask HN: DALL-E was trained on watermarked stock images?
#212Earlier quoted context omitted.
For text (the GPT-3 case), that’d work to train a model that had no knowledge of the last century of popular culture or idiom, and was significantly biased to more formal and traditional writing styles. The effects of this would be really quite interesting, but I think it would significantly limit the places it could be usefully applied. For DALL·E and Copilot, I’m confident that you couldn’t find anywhere near enoug…
IANAL, but what you should be able to do is have a set of quotes of one or two sentences each from various sources (books, TV, movies, etc.) that have the modern word or idiom you are specifying. That should then be small enough to fall under fair use (like IMDB, wikiquote, etc. have quotes from films, and good reads, dictionaries, etc. have quotes from books/text), and be plenty enough to capture the meaning of the…
- trademarks are protected even when the content itself wouldn't be copyrightable; you can't sell AI-generated T-shirts that "happen to" include the word Nike.
- NBC has a trademark on three tones, total length under 2 seconds[2]
[1]: https://fairuse.stanford.edu/2003/09/09/copyright_protection...
Re: Ask HN: DALL-E was trained on watermarked stock images?
#213All large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c... . When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Su…
It's not going to fail: the US courts are big company biased, and all the big companies are going to show out in force and money to ensure they get the result they want. But even extending that: knocking copyright'd images out isn't going to stop these systems. We know they work now, so if you have to be careful about licensing then that's just going to be done. The idea that any of these platforms will "die" if copy…
Knowing Disney, those terms would 100% include "subject to Disney's approval obtained prior to publication", and a chunk of money extracted from you. That company is very controlling of their IP.
Either that, or you're talking about a dataset that generates very specific looking images, that do not remind you of any major classic Disney property (so not "every piece"). Your own Jedi / Avenger avatar creator maybe, custom Mickey Mouse world character absolutely not.
Re: Ask HN: DALL-E was trained on watermarked stock images?
#214I am not a lawyer, but I've had to argue about copyright with several. In the United States, there are two bits of case law that are widely cited and relevant: In Kelly v. Arriba Soft Corp (9th), found that making thumbnails of images for use in a search engine was sufficiently "transformative" that it was ok. Another case, Perfect 10 (9th), found that thumbnails for image search and cached pages were also transforma…
It seems it is possible to generate images which are very similar to the existing stock photos if you feed getty images' description into DALL-E. I tried it with a distinctive banana image: https://imgur.com/a/0OrIr6e
You can't copyright an idea.
Re: Ask HN: DALL-E was trained on watermarked stock images?
#215Earlier quoted context omitted.
Nah, you could zap the training sets tomorrow and start over with public domain material and it would be fine. In fact I think you could easily get paid to generate more content for it.
For text (the GPT-3 case), that’d work to train a model that had no knowledge of the last century of popular culture or idiom, and was significantly biased to more formal and traditional writing styles. The effects of this would be really quite interesting, but I think it would significantly limit the places it could be usefully applied. For DALL·E and Copilot, I’m confident that you couldn’t find anywhere near enoug…
Re: Ask HN: DALL-E was trained on watermarked stock images?
#216Earlier quoted context omitted.
Perhaps I’m misunderstanding your argument, but my counterexample would be: if a human digital artist transformed a Getty image, resulting a fantastical, never-before-seen result, using software like Photoshop, that use would be no more defensible. If anything, the vast scale at which this occurs in AI makes it worse.
I think your hypothetical would depend on the character & extent of the transformation. Mere filters that leave the original recognizable? Probably an infringement. But creative application of transformations to express new ideas? Maybe not – especially if the derivative is a comment/parody on the original, that actually increases interest in it. Most art is a conversation with the past, reusing recognizable motifs &…
I find legal disputes in fine art interesting, however—IANAL, of course—I understand that fine artists (Richard Prince comes to mind) are subject to very different copyright restrictions than graphic artists under commercial use.
It’s, as you said, up to courts to decide. But AI generated imagery is frequently commercial in nature (KFC, already). AI services are trained on unlicensed commercial stock images, and are able to reproduce enormous quantities of derivative images, and do so at a profit. I think that’s categorically different from a fine artist appropriating imagery in a single artwork or even series of artworks in an entirely different context.
Re: Ask HN: DALL-E was trained on watermarked stock images?
#217Earlier quoted context omitted.
Top 1% is a bit exaggerated, but there is definitely a lot of not good stuff. I find that Dall-E does especially poorly with underspecified prompts too, unlike something like Midjourney which can give visually pleasing photos for even the most abstract concepts. Dall-E tends to do better with concrete and specific prompts. Here's an example: Stressful Shapes Dall-E: https://i.imgur.com/JBkSh0y.png Midjourney: https:/…
> On the other hand, here's a specific prompt: "nerdy yellow duck reading a magical book full of spells" > Dall-E: https://i.imgur.com/FMKZ8zc.png How well it learned all the common prejudices! "nerdy" == wears glasses I'm applauding. I'm looking already forward to AGI based on the current approaches… It will lead us finally into a better world, for sure. /s
Re: Ask HN: DALL-E was trained on watermarked stock images?
#218Earlier quoted context omitted.
But still "king of belgium giving a speech to an audience, but the audience members are cucumbers" is very specific. And I don't see the king of Belgium anywhere, two pictures have absolutely nothing to do with the prompt (no king, no speech, no audience, no cucumber), one has the speech and audience but no king or cucumber. Graphically, they are deep into the uncanny valley. Only the third image is kind of right, if…
I’ve been comparing Dall-E, MidJourney, and StableDiffusion. Goes to show how much training set and implementation choices matter. But in all cases, you have to think of the underlying labeled text-to-image sets as paint colors to mix, and prepare a palette accordingly. Still haven’t figured out how to get what I want, but to your point, one can get closer. - - - Not sure if this is why, but with OpenAI’s Dall-E, you…
Very insightful tip on how to harness the "creativity" of Dall-E and the like.
I see how the phrase "king of belgium" was too vague for Dall-E, so it didn't produce anything recognizable - but changing the words into known details, like "banker" and "salt and pepper hair", worked effectively to generate concrete imagery.
Hilarious results. :)
Re: Ask HN: DALL-E was trained on watermarked stock images?
#219Earlier quoted context omitted.
The synergistic effect of all the AI's inputs absolutely results in a unique new 'flair', with extensions, reversals, and mash-ups of styles just as in human-made artistic styles. And AI "builds exclusively on past experience and work of humans" just like any young new human artist equally does. In many cases, you can even tell the different models' outputs apart, not by raw quality or glitches, but by hard-to-descri…
Indeed the genie is out. And while we will get some interesting AI uses ultimately this is degenerative tech. In the end we end up with less authentic, less unpredictible and less delightfull art. Instead we get the perfectly suited to us, predictible, mediocre stuff. I said it in comment above - yes people build on work of others but they also bring lots of their originality and intelect. Part of what people do is t…
The thinking is still done by the human prompter.
Re: Ask HN: DALL-E was trained on watermarked stock images?
#220Earlier quoted context omitted.
I tried it with Stable Diffusion as well. You can use actual people and the model is even pretty decent at many of the famous ones. On the other hand, it is more difficult to get it to produce absurd results like these. my prompt: King Philippe I. of Belgium giving a speech surrounded by [[[[large green vertical cucumbers]]]], digital art in the style of Greg Rutkowski https://files.catbox.moe/1ej1a4.png
Is the usage of the surrounding brackets some kind of keyword weights specific to stable diffusion?