Live data from Hacker News

DALL·E now available in beta

openai.com

411–420 of 579 posts

Re: DALL·E now available in beta

#412

I have been having a blast with DALL-E, spending about an hour a day trying out wild combinations and cracking my friends up. I cannot imagine getting bored of it; it's like getting bored with visual stimulus, or art in general. In fact, I've been glad to have a 50/day limit, because it helps me contain my hyperfocus instincts. The information about new pricing is, to me as someone just enjoying making crazy imagines…

Yeah I've been having fun with it recreating bad Heavy Metal album art (https://twitter.com/P_Galbraith/status/1548597455138463744). It's good, but surprisingly difficult to direct it when you have a composition in mind. A few of these I burned through 20-30 prompts to get and I can't see myself forking up hundreds of dollars to roll the dice.

My brother is a digital artist and while excited at first he found it to be not all that useful. Mainly because it falls apart with complex prompts, especially when you have a few people or objects in a scene, or specific details you need represented, or a specific composition. You can do a lot with in-painting but it requires burning a lot of credits.

Re: DALL·E now available in beta

#413

Earlier quoted context omitted.

I think people don't realize how huge these models really are. When they're free, it's pretty cool. But charge an amount where there's actual profit in the product? Suddenly seems very expensive and not economically viable for a lot of use cases. We are still in the "you need a supercomputer" phase of these models for now. Something like DALLE mini is much more accessible but the results aren't good enough. Early ear…

> I think people don't realize how huge these models really are. They really aren't that large by the contemporary scaling race standards. DALLE-2 has 3.5B parameters, which should fit on an old GPU like Nvidia RTX2080, especially if you optimize your model for inference [1][2] which is commonly done by ML engineers to minimize costs. With optimized model, your memory footprint is ~1 byte per parameter, and some less…

It's true that image models are much less of a burden on GPU VRAM than a model like BLOOM where fitting it into a few A100s is ideal, but these diffusion models are a PITA for a ordinary hobbyist in terms of total compute: the CLIP pass over the text input is almost free, but then you feed it into the diffusion model, for one sample you'll be doing 10-100 forward passes (depending on how fancy the diffusion methods are - maybe even 1000 passes if you're using older/simpler ones), and for interactive use, you really want more like 6-9 separate samples; then they have to pass through the upscalers, which are diffusion models themselves and need to do a bunch of forward passes to denoise it. If you do 1 sample in 10s on your 1 consumer GPU, which would be pretty good, 6-9 means a joykilling minute+ wait. And then the user will pick a variation or edit one or tweak the prompt, and start all over again! It's like being back on 16kb dialup waiting for the newest .com to load.

Re: DALL·E now available in beta

#415

Since many people will start generating their first images soon, be sure to check out this amazing DALL-E prompt engineering book [0]. It will help you get the most out of DALL-E. [0]: https://dallery.gallery/wp-content/uploads/2022/07/The-DALL%... (PDF)

AMAZING, Thank you.

I hope that every science teacher that can - provide this to every student. This is the future they live in now. They should know these as well as they know how to install an app on a device.

Wait until we have a DALL-E -- Enabled Custom EMOJI stream - whereby, every text you send out has it corresponding DALL-E resultant image for every txt --

Then we can compare images from different people at different times but the prompt was identical... and see what the resultant library of emojiPROMPT looks like?

What about using Dall-e as a watermark for 'nft' signature 'notary' of an email.

If DALL-E provided a unique PID# for every image - and that PID was a key that only the OP runner of the image has - it can be used to authenticate an image to a text source... ??? (Assuming that no two prompts have the same result ever, but assigning a unique id that CAN be used to replay the image to verify it was generated when an original email/SMS was actually sent - it could be a unique way to timestamp authenticity/provenance of a thing...

Re: DALL·E now available in beta

#416

Earlier quoted context omitted.

I'm sure the novelty wears off. But I'm already coming up with several applications for it. On the personal side, I've been getting into game development, but the biggest roadblock is creating concept art. I'm an artist but it takes a huge amount of time to get the ideas on paper. Using DALLE will be a massive benefit and will let me expedite that process. It's important to note that this is not replacing my entire c…

I did notice it is very good at making small pixel art icons/sprites.

Do you have any tips or some prompt examples?

Re: DALL·E now available in beta

#417
post #337

Wait until someone trains a model like this, for porn. There seems to be a post-DALLE obscenity detector on openAI's tool, as so far I've found it to be entirely robust against deliberate typos designed to avoid simple 'bad word lists'. Ask it for a "pruple violon" and you get purple violins... you get the deal. "Metastable" prompts that may or may not generate obscene (content with nudity, guns, violence as I've fou…

I’ve thought about this and in fact porn generation sounds like a good thing?? It ensures that it’s victimless. Of course, there is a problem with generation of illegal (underage) porn but other than this, I think it could be helpful for this world.

i wonder if you could say “person who looks 15 but is 18”

Re: DALL·E now available in beta

#418

The name "OpenAI" to me implies being open-source. I have an RTX 3080 and will likely be buying a 4090 when it comes out. Will I ever be able to generate these images locally, rather than having to use a paid service? I've done it with DALL-E Mini, but the images from that don't hold a candle to what DALL-E 2 produces.

You should show up to the US Open with a tennis racket next year and see if they'll let you have a go, too.

Re: DALL·E now available in beta

#419
post #58

I was supposed to be making a video game, but got a bit sidetracked when DALL·E came out and made this website on the side: http://dailywrong.com/ (yes I should get SSL). It's like The Onion, but all the articles are made with GPT-3 and DALL·E. I start with an interesting DALL·E image, then describe it to GPT-3 and ask it for an Onion-like article on the topic. The results are surprisingly good.

It would be interesting to see if there was a market for a monthly newsprint version similar to the old https://weeklyworldnews.com/

Can DALL-E render Bat Boy?

Re: DALL·E now available in beta

#420
post #396

Earlier quoted context omitted.

Dall-e has novelty, but no intent, meaning, originality. Yes the author can be creative at generating prompts, but visually I haven’t seen it generate anything that feels artistically interesting. If you want pre-existing concepts in novel combinations then yes it works. It’s good at “in the style of” but there’s no “in a new style”. It has a house style too that tends to feel Reddit-like.

Isn't every "new style" just a novel combination of pre-existing concepts? Nothing new under the sun and all that. Either way, I feel like your view is an exhaustingly pessimistic take on AI-generated art. I mean, sure, most of what DALL-E generates is pretty mundane, but other times I have been surprised at how bizarre and unique certain images are. You seem to imply that because an AI is not human, its art is not i…

> Isn't every "new style" just a novel combination of pre-existing concepts?

At the extreme limit, maybe. But within art or even digital art, then new styles are actually not that rare, humans are pretty good at generating them. Maybe they grab inspiration from nature, visual phenomena, etc, so in that sense it's not "new" but it is "new to the medium". In art you new styles all the time. DALL-E will never do that by it's very nature, and so it's easy to see how it's boring.

And that's just the stylistic level, but it's happening at almost all levels. It's almost definitional that it doesn't innovate, only remix.

It's strange framing this as pessimistic, it's not really optimistic nor pessimistic, it just is. It's also not AI, and that's important to realize: it's a statistical model that generates purely based on pre-existing training. It's very nature is without-meaning and without-originality. That doesn't detract from it being cool or interesting or helpful or enjoyable. I find it cool and useful.

But it's not innovative or creative or meaningful by itself.

Post reply on HN