Live data from Hacker News

Dall-E 2

openai.com

391–400 of 511 posts

Re: Dall-E 2

#391

Apologies for an open-ended question but: does anyone know if there is a term for something like Turing-completeness within AI, where a certain level of intelligence can simulate any other type of intelligence like our brains do? For example, using DeMorgan's theorem, we can build any logic circuit out of all NAND or NOR gates: https://www.electronics-tutorials.ws/boolean/demorgan.html https://en.wikipedia.org/wiki/N…

> Apologies for an open-ended question but: does anyone know if there is a term for something like Turing-completeness within AI, where a certain level of intelligence can simulate any other type of intelligence like our brains do?

> So how can we adapt these AI experiments to get real work done?

You're missing a step here - the difference between "imagining doing something" and "actually doing something". An ML model can produce thoughts, but that isn't necessarily the same direction of research as actually doing things in real life, much less becoming superhuman and taking over the world etc.

In your imagination, everything always goes your way.

Re: Dall-E 2

#392
post #73
post #59

I'm only part way through the paper, but what struck me as interesting so far is this: In other text-to-image algorithms I'm familiar with (the ones you'll typically see passed around as colab notebooks that people post outputs from on Twitter), the basic idea is to encode the text, and then try to make an image that maximally matches that text encoding. But this maximization often leads to artifacts - if you ask for…

What always bother me with this stuff is, well, you say one approach is more sensible than the other because the images happen to come out more pleasing. But there's no real rhyme or reason, it is a sort of alchemy. Is text encoding strictly worse or is it an artifact of the implementation? And if it is strictly worse, which is probably the case, why specifically? What is actually going on here? I can't argue that th…

> But there's no real rhyme or reason, it is a sort of alchemy.

Is there a rhyme or reason as to why picasso decided to paint like that? Yes these networks are hard to reason about, but so are real human brains.

Re: Dall-E 2

#393

Earlier quoted context omitted.

Thanks! Do you happen to know how much GPU RAM I need to run glid-3 and/or the latent diffusion model, if I don't want to run on colab?

Just tried glid-3 with a batch size of one and I'm getting 4781MiB. The latent diffusion model peaks at 8403MiB These are fp16 numbers though, you might need a recent nvidia card to run it.

I'll try them out. I have an RTX 2070, which apparently supports fp16. But it only has 8GB RAM.

I used the instructions here to check: https://github.com/wang-xinyu/tensorrtx/blob/master/tutorial...

Re: Dall-E 2

#394

Earlier quoted context omitted.

That's not how this works. There is no 'search' step, there is no 'superimposing' step. It's not really possible to explain what the AI is doing using these concepts. If you pay attention to all the corgi examples, the sofa texture changes in each of them, and it synthesizes shadows in the right orientation - that's what it's trained to do. The first one actually does give you the impression of weight. And if you loo…

How do you propose we talk about what it is doing if not by using the terminology from the human editing process it is replacing? I'm struggling to express things. My issue is that it appears to not be possible to explain what the AI is doing at all. If you could, you'd be able to actually control the output. And talking about how the model is trained is interesting but not an answer. Of course there is a superimposi…

Being opaque to human understanding is one of the downsides of existing AI/ML tech, for sure. Check out to the video in the page, and notice how the images transition from random color blobs to increasing detail - that's showing you how the image is being generated. It's a continuous process of trying to satisfy a prediction, there are no discrete editing steps.

The kind of tech you're imagining, where the computer has semantic understanding of what's in the picture, and is reproducing something based on a 3D scene, knowledge of physics, materials, etc is probably decades away. In that sense yes, this is just a 'trick'.

Re: Dall-E 2

#395

Earlier quoted context omitted.

That's not how this works. There is no 'search' step, there is no 'superimposing' step. It's not really possible to explain what the AI is doing using these concepts. If you pay attention to all the corgi examples, the sofa texture changes in each of them, and it synthesizes shadows in the right orientation - that's what it's trained to do. The first one actually does give you the impression of weight. And if you loo…

How do you propose we talk about what it is doing if not by using the terminology from the human editing process it is replacing? I'm struggling to express things. My issue is that it appears to not be possible to explain what the AI is doing at all. If you could, you'd be able to actually control the output. And talking about how the model is trained is interesting but not an answer. Of course there is a superimposi…

> But Vermeer could explain how he came up with the style, his techniques, choices, 'etc.

Often they can't. Ramanujan couldn't explain how he solved math problems, for instance, and humans can forget their own history easily, or even forget how to do something consciously while still doing it through muscle memory.

An ML model wouldn't forget the same way, but it could just lie to you.

Re: Dall-E 2

#397

Earlier quoted context omitted.

Regarding cherry-picking, the images of astronauts on horses look stunning, except for their hands. There's something seriously wrong with their hands. Maybe give it another five years, a few more $billion and a few more petabytes/flops and it will be good. Then finally everyone can generate art for their own Magic: the Gathering cards. (That's the end goal, right?)

Interestingly, hands are also something humans struggle to draw. They're a very complex anatomical form, many small tendons and muscles. Many artists struggle to depict hands. They're not made out of a few straight lines like a torso, there's lots of skew going on. They're probably the hardest structure of the human body to 'learn' for a ML system.

I think some of this is because hands are very involved in both communication and threat assessment, so we as humans put a lot of automatic attention on them. We aren't even usually aware of it--unless something looks off

Personally I really like leaning into the discontinuities and quirkiness of generated images. This is output from GLID-3: https://twitter.com/mwegner/status/1511139661095178241

Re: Dall-E 2

#398
post #353

>We’ve limited the ability for DALL·E 2 to generate ... adult images. I think that using something like this for porn could potentially offer the biggest benefit to society. So much has been said about how this industry exploits young and vulnerable models. Cheap autogenerated images (and in the future videos) would pretty much remove the demand for human models and eliminate the related suffering, no? EDIT: typo

No. If people are exposed to stimuli, they will pursue increasingly stimulating versions of it. I.e., if they see artificial CP, they will often begin to become desensitized (habituated) and pursue real CP or even live children thereafter. Conversely, if people are not exposed to certain stimuli, they will never be able to conceptualize them, and thus will be unable to think about them. Obviously you cannot eliminate…

I'm not sure I agree with the statement, you're putting forth a lot of assertions without the actual quantitative data to back up what you're saying, and even though you think it sounds intuitive that doesn't necessarily make it valid.

I'd actually argue the reverse, I think you see a lot more effort towards acquiring things that are illegal than you would otherwise.

Re: Dall-E 2

#399

Earlier quoted context omitted.

The tail end of programming will be the last thing to be replaced, maybe. I don’t see why CRUD apps get to hide under the umbrella of programming ultra-advanced AI.

Let me know when you can speak English to a computer and have it generate CRUD code that satisfies all engineering and design constraints. The AI will need to be dynamic enough to understand nuance, missing gaps in the requirements spec, have context on the application being built, able to suggest improvements on product design, know how to make changes through the same conversational interface, etc. Accomplishing th…

Let me know when you can speak English to a computer and have it generate CRUD code that satisfies all engineering and design constraints. The AI will need to be dynamic enough to understand nuance, missing gaps in the requirements spec, have context on the application being built, able to suggest improvements on product design, know how to make changes through the same conversational interface, etc.

Let me know when you find a single programmer who can do that reliably.

Re: Dall-E 2

#400

Earlier quoted context omitted.

I have degrees and several years of experience in both fields, and I can tell you that both are creative professions where output is unbounded and the measure of success is subjective; these are the fields that will be safe for a while. IMO it's fields such as aircraft pilots who should be most worried.

The jobs of commercial pilots are very safe. Pilots are not there to fly the aircraft, the autopilot already does that. They are there to command the aircraft, in a pair in case one is incapacitated, making the best decisions for the people on board, and to troubleshoot issues when the worst happens. No AI or remote pilot is going to help when say... the aircraft loses all power. Or the airport has been taken over in…

It's not as safe as you believe it to be, in the case of total electrical power failure in a fly by wire airliner, and the corresponding loss of hydraulic pressure there's very little that a pilot can do with that point.

As far as the extremely unlikely hostage situation goes, if it were AI controlled that would be even less likely attempts from people to hijack an airplane in the first place since there wouldn't be a human element a.k.a. a pilot that they could appeal to their emotion.

Post reply on HN