Live data from Hacker News

DALL·E now available in beta

openai.com

271–280 of 579 posts

Re: DALL·E now available in beta

#271

Earlier quoted context omitted.

I think people don't realize how huge these models really are. When they're free, it's pretty cool. But charge an amount where there's actual profit in the product? Suddenly seems very expensive and not economically viable for a lot of use cases. We are still in the "you need a supercomputer" phase of these models for now. Something like DALLE mini is much more accessible but the results aren't good enough. Early ear…

What are the resources at work here? What are the resources needed to train this model? If someone just gave you the model for free, what resources would you need to use it to generate new results?

If I had to guess, based on other large models, it’s in the range of hundreds of GBs. It might even be in the TB range. To host that model for fast production SaaS inference requires many GPUs. An A100 has 80GB, so a dozen A100s just to keep it in memory, and more if that doesn’t meet the request demand.

Training requires even more GPUs, and I wouldn’t be surprised if they used more than 100 and trained over 3 months.

Re: DALL·E now available in beta

#273

I have been having a blast with DALL-E, spending about an hour a day trying out wild combinations and cracking my friends up. I cannot imagine getting bored of it; it's like getting bored with visual stimulus, or art in general. In fact, I've been glad to have a 50/day limit, because it helps me contain my hyperfocus instincts. The information about new pricing is, to me as someone just enjoying making crazy imagines…

Check out Artbreeder, it is likewise a ton of fun!

Multimodal.art (https://multimodal.art/) is working on a free version of something like DALLE, though it's not that good as of yet.

Re: DALL·E now available in beta

#274
post #37

Earlier quoted context omitted.

AFAIK only people can own copyright (the monkey selfie case tested this), and machine-generated outputs don't count as creative work (you can't write an algorithm that generates every permutation of notes and claim you own every song[1]), so DALL-E-generated images are most likely copyright-free. I presume OpenAI only relies on terms of service to dictate what users are allowed to do, but they can't own the images, a…

The monkey selfie was not derived from millions of existing works, and that is the difference. If an artist has a well-known art style, and this algorithm was trained on it and can copy that style, would the artist have grounds to sue? I don't know.

Even if you imitate someone's style intentionally, they don't have grounds to sue. Style isn't copyrightable in the US. Whether DALL-E outputs are a derivative work is a different question, though

Re: DALL·E now available in beta

#275

Earlier quoted context omitted.

If I was using this in South Korea, how is showing all white people any better than showing whites, blacks, latinos and asians?

You would presumably input “South Korean CEO”. DALL-E would then unhelpfully add “black” “female” without your knowledge.

I just tried it out and it looks like DALL-E isn't as inept as you imagined. Exact query used was 'A profile photo of a male south korean CEO', and it spat out 4 very believable korean business dudes.

Supplying the race and sex information seems to prevent new keywords from being injected. I see no problem with the system generating female CEOs when the gender information is omitted, unless you think there are?

Re: DALL·E now available in beta

#276
post #259

Earlier quoted context omitted.

I've only seen this thing https://huggingface.co/spaces/dalle-mini/dalle-mini is it not dall-e?

It is not and that's why OpenAI asked them to change the name, which they did.

oh. I retract my OP then

Re: DALL·E now available in beta

#277
post #196

Earlier quoted context omitted.

This is clever. Does GPT-3 come up with the title of the article, too? That's the funniest part.

At first I came up with them myself, but found that it often comes up with better ones, so I ask it for variations. I think I got it to even fill the title given a picture, something like “Article picture caption: Man holding an apple. Article title: ...”. Might experiment more with that in the future.

Well, then I'm impressed with GPT-3's ability to generate those titles!

The combination of photo/title feels like they come from the more absurd articles published by theonion.

If we aren't living in a simulation, it's just a matter of time...

Re: DALL·E now available in beta

#278

Earlier quoted context omitted.

I have tried playing around with the beta access to make it generate NFT art with different prompts, but in vail. I think it has not been trained on NFT art (crypto punks and so on).

Heads up: I think you meant "in vain" rather than "in vail". However, a similar phrase is "to no avail" which also means that something was not successful.

I think you meant "in vain" rather than "in vein".

Re: DALL·E now available in beta

#279

Earlier quoted context omitted.

You would presumably input “South Korean CEO”. DALL-E would then unhelpfully add “black” “female” without your knowledge.

I just tried it out and it looks like DALL-E isn't as inept as you imagined. Exact query used was 'A profile photo of a male south korean CEO', and it spat out 4 very believable korean business dudes. Supplying the race and sex information seems to prevent new keywords from being injected. I see no problem with the system generating female CEOs when the gender information is omitted, unless you think there are?

Isn’t the diversity keyword injection random?

My point is that it is pointless. If you want an image of a person included, you can just specify it yourself.

Re: DALL·E now available in beta

#280

Earlier quoted context omitted.

Your input isn't being polluted by this any more than it is when the tokens in it are ground up into vectors and transformed mathematically. You just have an easier time understanding this transformation.

Obviously, it's polluted. Undisputably. In a mathematical sense, an extra (black box) transformation is performed on the input to the model. In a practical sense (eg. if you're researching the model), this is like having dirty laboratory tools - all measurements are slightly off. The presumption by OpenAI is that the measurements are off in the correct way . I'm interested in using Dall-E commercially, but I think so…

It's a fucking AI picture generator. The whole thing is a series of (literally) inscrutable black boxes. This is not a good argument.
Post reply on HN