Live data from Hacker News

Dall-E 2

openai.com

201–210 of 511 posts

Re: Dall-E 2

#202

A few comments by someone who's spent way too much time in the AI-generated space: * I recommend reading the Risks and Limitations section that came with it because it's very through: https://github.com/openai/dalle-2-preview/blob/main/system-c... * Unlike GPT-3, my read of this announcement is that OpenAI does not intend to commercialize it, and that access to the waitlist is indeed more for testing its limits (and…

Not-so-open.ai

open-your-wallet.ai

Re: Dall-E 2

#203
Is there a geometric model relative to this? EG: "corgi near the fireplace" but the output is a 3d model of the corgi and fireplace with shaders rather than an image.

Re: Dall-E 2

#204
Dall-E 2 seems to be incapable to catch the essence of the art. I'm not really surprised by it, I'd be surprised a lot if it could. But nevertheless: if you looked in the eye of a Girl With A Pearl Earring[1], you'd be forced to stop and to think what does she have on her mind right now. Or may be you had some other question in your mind, but it really stops people to think. But none of Dall-E interpretations have this quality. Works inspired by Girl With A Pearl Earring sometimes have at least part of that power, like Girl With a Babmoo Earring[2]. But none of Dall-E interpretations have such a power.

And this observation may lead to a great consequences for visual arts. I had a lot of joy of looking at different Dall-E interpretations to find what the flaw of the interpretation that forbids it to be a piece of art of an equal value to the original. It is a ready made tool to search for explanations of the Power of Art. It cannot say what detail make a picture to be an artwork, but it allow to see multiple data points, and to narrow the hypothesis space. My main conclusion is that the pearl earring have nothing to do with the power of art. It is something in the eye, and probably with the slightly opened mouth. (Somehow Dall-E pictured all interpretations with closed lips, so it seems to be an important thing, but I need more variation along this axis to be sure).

[1] https://en.wikipedia.org/wiki/Girl_with_a_Pearl_Earring [2] https://yourartshop-noldenh.com/awol-erizku-girl-with-the-pe...

Re: Dall-E 2

#205

It's becoming clear that efficient work in the future will hinge upon one's ability to accurately describe what one wants . Unpacking that -- a large piece is the ability to understand all the possible "pitfalls" and "misunderstandings" that could happen on the way to a shared understanding. While technical work will always have a place -- I think that much creative work will become more like the management of a team…

> accurately describe what one wants

Isn't that essentially what programming already is?

Re: Dall-E 2

#207

Earlier quoted context omitted.

> It is compositing as final step. I might be misinterpeting your use of "compositing" here (and my own technical knowledge is fairly shallow) but I don't think there's any compositing of elements generally in AI image generation. (unless Dall-E 2 changes this. I haven't read the paper yet)

https://cdn.openai.com/papers/dall-e-2.pdf > Given an image x, we can obtain its CLIP image embedding zi and then use our decoder to “invert” zi, producing new images that we call variations of our input. .. It is also possible to combine two images for variations. To do so, we perform spherical interpolation of their CLIP embeddings zi and zj to obtain intermediate zθ = slerp(zi, zj , θ), and produce variations of z…

I'm not sure that's "compositing" except in the most abstract sense? But maybe that's the sense in which you mean it.

I'd argue that at no point is there a representation of a "teddy bear" and "a background" that map closely to their visual representation - that are combined.

(I'm aware I'm being imprecise so give me some leeway here)

Re: Dall-E 2

#209

It's becoming clear that efficient work in the future will hinge upon one's ability to accurately describe what one wants . Unpacking that -- a large piece is the ability to understand all the possible "pitfalls" and "misunderstandings" that could happen on the way to a shared understanding. While technical work will always have a place -- I think that much creative work will become more like the management of a team…

Programming, art, music, is just “describing what you want” in a very specific way. This is describing what you want in a much more vague way.

The upside it that it’s more “intuitive” and requires much less detail and technique, as the AI infers the detail and technique. The downside is that it’s really hard to know what the AI will generate or get it to generate something really specific.

I believe the future will combine the heuristics of AI-generation with the specificity of traditional techniques. For example, artists may start with a rough outline of whatever they want to draw as a blob of colors (like in some AI image-generation papers). Then they can fill in details using AI prompts, but targeting localized regions/changes and adding constraints, shifting the image until it’s almost exactly what they imagined in their head.

Re: Dall-E 2

#210
post #144

Earlier quoted context omitted.

I mean was he really wrong? As models like OpenAI Codex get more powerful over time, they will start eating into large chunks of dev work as well...

Literally everyone on this website is in denial. They all approach it by asking which fields will be safe. No field is safe. “But it’s not going to happen for a long time.” Climate deniers say the same thing and you think they should be wearing the dunce hat? The average person complains bitterly about climate deniers who say that it’s “my grandkids problem lol” but when I corner the average person into admitting AI…

> but when it comes to the biggest threat to human well-being in history

Evolution doesn't stop for anyone, don't think like a dinosaur.

Post reply on HN