Live data from Hacker News

Dall-E 2

openai.com

121–130 of 511 posts

Re: Dall-E 2

#121
post #90
post #73

Earlier quoted context omitted.

What always bother me with this stuff is, well, you say one approach is more sensible than the other because the images happen to come out more pleasing. But there's no real rhyme or reason, it is a sort of alchemy. Is text encoding strictly worse or is it an artifact of the implementation? And if it is strictly worse, which is probably the case, why specifically? What is actually going on here? I can't argue that th…

Yeah, I mean you're right that ultimately the proof is in the pudding. But I do think we could have guessed that this sort of approach would be better (at least at a high level - I'm not claiming I could have predicted all the technical details!). The previous approaches were sort of the best that people could do without access to the training data and resources - you had a pretrained CLIP encoder that could tell you…

[deleted]

Re: Dall-E 2

#122
A few comments by someone who's spent way too much time in the AI-generated space:

* I recommend reading the Risks and Limitations section that came with it because it's very through: https://github.com/openai/dalle-2-preview/blob/main/system-c...

* Unlike GPT-3, my read of this announcement is that OpenAI does not intend to commercialize it, and that access to the waitlist is indeed more for testing its limits (and as noted, commercializing it would make it much more likely lead to interesting legal precedent). Per the docs, access is very explicitly limited: (https://github.com/openai/dalle-2-preview/blob/main/system-c... )

* A few months ago, OpenAI released GLIDE ( https://github.com/openai/glide-text2im ) which uses a similar approach to AI image generation, but suspiciously never received a fun blog post like this one. The reason for that in retrospect may be "because we made it obsolete."

* The images in the announcement are still cherry-picked, which is therefore a good reason why they tested DALL-E 1 vs. DALL-E 2 presumably on non-cherrypicked images.

* Cherry-picking is relevant because AI image generation is still slow unless you do real shenanigans that likely compromise image quality, although OpenAI has likely a better infra to handle large models as they have demonstrated with GPT-3.

* It appears DALL-E 2 has a fun endpoint that links back to the site for examples with attribution: https://labs.openai.com/s/Zq9SB6vyUid9FGcoJ8slucTu

Re: Dall-E 2

#123
post #103
post #97

Earlier quoted context omitted.

I don't think it is actually painting at all but I need to read the paper carefully. I think it is using a free text query to select the best possible clipart from a big library and blends it together. Still very interesting and useful. It would be extremely impressive if the "Kuala dunking a basketball" had a puddle on the court in which it was reflected correctly, that would be mind blowing.

This is actual image generation - the 'decoder' takes as input a latent code (representing the encoding of the text query), and synthesizes an image. It's not compositing or querying a reference library. The only time that real images enter the process is during training - after that, it's just the network weights.

It is compositing as final step. I understand that the Kuala it is compositing may have been a previously un-existent Kuala that it synthesized from a library of previously tagged Kuala images... that's cool, but what is the difference really from just plucking one of the pre-existing Kualas into the scene?

The difference is just that it makes the compositing easier. If you don't have a pre-existing image that would match the shadows and angles you can hallucinate a new Kuala that does. Neat trick.

But I bet if I threw the poor marsupial at a basket net it would look really differently than the original clipart of it climbing some tree in a slow and relaxed manner. See what I mean?

Maybe Dall-E 2 can make it strike a new pose. The limb positions could be altered. But the facial expression?

And if the basketball background has wind blowing leaves in one direction the Kuala fur won't match, it will look like the training set fur. The puddle won't reflect it. 'etc.

This thing doesn't understand what a Kuala is like a 3-yr old. It understands the text "Kuala" is associated with that tagged collection of pixel blobs and can conjure up similar blobs unto new backgrounds - but it can't paint me a new type of Kuala that it hasn't seen before. It just looks that way.

Re: Dall-E 2

#124
post #3

Preventing Harmful Generations We’ve limited the ability for DALL·E 2 to generate violent, hate, or adult images. By removing the most explicit content from the training data, we minimized DALL·E 2’s exposure to these concepts. We also used advanced techniques to prevent photorealistic generations of real individuals’ faces, including those of public figures. "And we've also closed off a huge range of potentially int…

This AI is still a minor. It can start looking at R rated images when it turns 17.

This is an apt analogy -- ensure that the model is mature enough to handle mature content.

Re: Dall-E 2

#125
post #3

Preventing Harmful Generations We’ve limited the ability for DALL·E 2 to generate violent, hate, or adult images. By removing the most explicit content from the training data, we minimized DALL·E 2’s exposure to these concepts. We also used advanced techniques to prevent photorealistic generations of real individuals’ faces, including those of public figures. "And we've also closed off a huge range of potentially int…

If you went to an artist who takes commissions and they said "Here are the guidelines around the commissions I take" would you complain in the same way? Who cares if it's a bunch of engineers or an artist. If they have boundaries on what they want to create, that's their prerogative.

To take that a step further, I wont code malware. I've never been asked but I'd refuse if I was. Everyone has their choices.

Re: Dall-E 2

#126
post #50

The correct response here from the artists point of view should be a widespread coming together against their art being used as training data for ML models. With a quickly spread new license on most major art submission sites that explicitly forbids AI algorithms from using their work, artists would effectively starve OpenAI and others from using their own works to put them out of a job.

The license should forbid competing artists to using the artist’s work as well. In fact, no human should come in contact with the produced art, otherwise they might be accidentally inspired by it, thus stealing from the original creator.

Re: Dall-E 2

#127
post #62

This reminds me of the holodeck in Star Trek. Someone could walk into the Holodeck and say “make a table in the center of the room. Make it look old.” It seemed amazing to me that the computer could make anything and customize it with voice. We are pretty close to star trek technology now in computer ability (ship’s computer, not Commander Data). I guess to really be like the holodeck it needs to be able to do 3d and…

[deleted]

Re: Dall-E 2

#129
Yeah, I mean you're right that ultimately the proof is in the pudding.

But I do think we could have guessed that this sort of approach would be better (at least at a high level - I'm not claiming I could have predicted all the technical details!). The previous approaches were sort of the best that people could do without access to the training data and resources - you had a pretrained CLIP encoder that could tell you how well a text caption and an image matched, and you had a pretrained image generator (GAN, diffusion model, whatever), and it was just a matter of trying to force the generator to output something that CLIP thought looked like the caption. You'd basically do gradient ascent to make the image look more and more and more like the text prompt (all the while trying to balance the need to still look like a realistic image). Just from an algorithm aesthetics perspective, it was very much a duct tape and chicken wire approach.

The analogy I would give is if you gave a three-year-old some paints, and they made an image and showed it to you, and you had to say, "this looks like a little like a sunset" or "this looks a lot like a sunset". They would keep going back and adjusting their painting, and you'd keep giving feedback, and eventually you'd get something that looks like a sunset. But it'd be better, if you could manage it, to just teach the three-year-old how to paint, rather than have this brute force process.

Obviously the real challenge here is "well how do you teach a three-year-old how to paint?" - and I think you're right that that question still has a lot of alchemy to it.

Re: Dall-E 2

#130
post #5

Some freely available models GLID-3: https://colab.research.google.com/drive/1x4p2PokZ3XznBn35Q5B... and a new Latent Diffusion notebook: https://colab.research.google.com/github/multimodalart/laten... have both appeared recently and are getting remarkably close to the original Dall-E (maybe better as I can't test the real thing...) So - this was pretty good timing if OpenAI want to appear to be ahead of the pack. Of…

I think this is really neat, but definitely not on the same tier as DALL-E 2, at least from the cherry-picked images I saw.
Post reply on HN