Live data from Hacker News

Dall-E 2

openai.com

241–250 of 511 posts

Re: Dall-E 2

#241
post #204

Dall-E 2 seems to be incapable to catch the essence of the art. I'm not really surprised by it, I'd be surprised a lot if it could. But nevertheless: if you looked in the eye of a Girl With A Pearl Earring[1], you'd be forced to stop and to think what does she have on her mind right now. Or may be you had some other question in your mind, but it really stops people to think. But none of Dall-E interpretations have th…

Initial Outputs from New AI Model Not As Good at Nuance as Historic Artwork, Approach Deemed Hopeless

Re: Dall-E 2

#242
post #235

Earlier quoted context omitted.

It is compositing as final step. I understand that the Kuala it is compositing may have been a previously un-existent Kuala that it synthesized from a library of previously tagged Kuala images... that's cool, but what is the difference really from just plucking one of the pre-existing Kualas into the scene? The difference is just that it makes the compositing easier. If you don't have a pre-existing image that would…

>And if the basketball background has wind blowing leaves in one direction the Kuala fur won't match, it will look like the training set fur. The puddle won't reflect it. If you read the article, it gives examples that do exactly this. For example, adding a flamingo shows the flamingo reflected in a pool. Adding a corgi at different locations in a photo of an art gallery shows it in picture style when it's added to a…

Well not so much an article as really interesting hand picked examples. The paper doesn't address this as far as I can tell. My guess is that this is a weak point that will trip it up occasionally.

A lot of the time it doesn't super matter, but sometimes it does.

Re: Dall-E 2

#243
post #61

Earlier quoted context omitted.

I mean was he really wrong? As models like OpenAI Codex get more powerful over time, they will start eating into large chunks of dev work as well...

Large chunks, yes, but all that means is that engineers will move up the abstraction stack and become more efficient, not that engineers will be replaced. Bytecode -> Assembly -> C -> higher level languages -> AI-assisted higher-level languages

History isn't a great guide here. Historically the abstractions that increased efficiency begat further complexity. Coding in Python elides over low-level issues but the complexity of how to arrange the primitives of python remains for the programmer to engage with. AI coding has the potential to elide over all the complexity that we identify as programming. I strongly suspect this time is different.

The space for "AI-assisted higher-level languages" sufficiently distinct from natural language is vanishingly small. Eventually you're just speaking natural language to the computer, which just about anyone can do (perhaps with some training).

Re: Dall-E 2

#244
post #238
post #73

Earlier quoted context omitted.

What always bother me with this stuff is, well, you say one approach is more sensible than the other because the images happen to come out more pleasing. But there's no real rhyme or reason, it is a sort of alchemy. Is text encoding strictly worse or is it an artifact of the implementation? And if it is strictly worse, which is probably the case, why specifically? What is actually going on here? I can't argue that th…

> But if doesn't spit up exactly what you want can't edit it further. Why? You can tweak the prompt, change parameters, or even use the actual "edit" capability that they demo in the post.

Maybe I am misunderstanding but if you start tweaking the prompt you'll end with something completely different.

The "edit" capability, as far as I can tell please correct me if I got confused, is picking your favorite out of the generated variations.

I would like to "lock" the scene and add instructions like "throw in a reflection".

Re: Dall-E 2

#245
post #5

Some freely available models GLID-3: https://colab.research.google.com/drive/1x4p2PokZ3XznBn35Q5B... and a new Latent Diffusion notebook: https://colab.research.google.com/github/multimodalart/laten... have both appeared recently and are getting remarkably close to the original Dall-E (maybe better as I can't test the real thing...) So - this was pretty good timing if OpenAI want to appear to be ahead of the pack. Of…

With glide I think we've reached something of a plateau in terms of architecture on the "text to image generator S curve". DALL-E-2 is a very similar architecture to glide and has some notable downsides (poorer language understanding)

glid-3 is a relatively small model trained by a single guy on his workstation (aka me) so it's not going to be as good. It's also not fully baked yet so ymmv, although it really depends on the prompt. The new latent diffusion model is really amazing though and is much closer to DALLE-2 for 256px images.

I think the open source community will rapidly catch up with Openai in the coming months. The data, code and compute are all there to train a model of similar size and quality.

Re: Dall-E 2

#246
>We’ve limited the ability for DALL·E 2 to generate ... adult images.

I think that using something like this for porn could potentially offer the biggest benefit to society. So much has been said about how this industry exploits young and vulnerable models. Cheap autogenerated images (and in the future videos) would pretty much remove the demand for human models and eliminate the related suffering, no?

EDIT: typo

Re: Dall-E 2

#247
post #73
post #59

I'm only part way through the paper, but what struck me as interesting so far is this: In other text-to-image algorithms I'm familiar with (the ones you'll typically see passed around as colab notebooks that people post outputs from on Twitter), the basic idea is to encode the text, and then try to make an image that maximally matches that text encoding. But this maximization often leads to artifacts - if you ask for…

What always bother me with this stuff is, well, you say one approach is more sensible than the other because the images happen to come out more pleasing. But there's no real rhyme or reason, it is a sort of alchemy. Is text encoding strictly worse or is it an artifact of the implementation? And if it is strictly worse, which is probably the case, why specifically? What is actually going on here? I can't argue that th…

I think deep learning is better thought of as "science" than "engineering." Right now we're in the stage of the Greeks and Arabs where we know "if we do this then that happens." It will be awhile before we have a coherent model of it, and I don't think we will ever solve all of its mysteries.

Re: Dall-E 2

#248
post #73
post #59

I'm only part way through the paper, but what struck me as interesting so far is this: In other text-to-image algorithms I'm familiar with (the ones you'll typically see passed around as colab notebooks that people post outputs from on Twitter), the basic idea is to encode the text, and then try to make an image that maximally matches that text encoding. But this maximization often leads to artifacts - if you ask for…

What always bother me with this stuff is, well, you say one approach is more sensible than the other because the images happen to come out more pleasing. But there's no real rhyme or reason, it is a sort of alchemy. Is text encoding strictly worse or is it an artifact of the implementation? And if it is strictly worse, which is probably the case, why specifically? What is actually going on here? I can't argue that th…

No post body was provided.

Re: Dall-E 2

#249
post #204

Dall-E 2 seems to be incapable to catch the essence of the art. I'm not really surprised by it, I'd be surprised a lot if it could. But nevertheless: if you looked in the eye of a Girl With A Pearl Earring[1], you'd be forced to stop and to think what does she have on her mind right now. Or may be you had some other question in your mind, but it really stops people to think. But none of Dall-E interpretations have th…

art criticism should be off topic here. This is more like chopping off the visual cortex and some association cortex from a brain and stimulating it. there is no person signaling to us, nor can we attribute any striking images that may come up to a person with agency.

But its like a giant database of decent clipart for anything we can imagine

Re: Dall-E 2

#250

>We’ve limited the ability for DALL·E 2 to generate ... adult images. I think that using something like this for porn could potentially offer the biggest benefit to society. So much has been said about how this industry exploits young and vulnerable models. Cheap autogenerated images (and in the future videos) would pretty much remove the demand for human models and eliminate the related suffering, no? EDIT: typo

The problem might be that people are simply lying. Their real reasons are religious/ideological, but they cite humanitarian concerns (which their own religious stigma is partly responsible for).
Post reply on HN