Live data from Hacker News

DALL·E now available in beta

openai.com

301–310 of 579 posts

Re: DALL·E now available in beta

#301

Earlier quoted context omitted.

It's a fucking AI picture generator. The whole thing is a series of (literally) inscrutable black boxes. This is not a good argument.

Yeah man, but literally the entire point of this AI picture generator is that it's, like, super accurate at rendering the prompt, and stuff. I don't understand the relevance of the black box's scrutability - I just want to play with the black box . I am interested in increasing my understanding of the black box, not of a trust-me-it's-great-our-intern-steve-made-it black box derivative.

You should make your own black boxes then. By all means, send your dollars to whatever service passes your purity test; I'm just saying that the idea that DALL-E is "polluting" your input is risible. It's already polluting your data at, like, a subatomic level, at dimensionalities it hadn't even occurred to you to consider, and at enormous scale.

Re: DALL·E now available in beta

#302
post #182
post #76

I fully expect stock image sites to be swamped by DALL-E generated images that match popular terms (e.g. "business person shaking hands"). Generate the image for $0.15. Sell it for $1.00.

DALLE images are still only 1024 px wide. Which has its uses, but I don’t think the stock photo industry is in real danger until someone figures out a better AI superresolution system that can produce larger and more detailed images.

I've been using this app to upscale the images to 4000x4000, and it works amazingly well (there is also a version for Android):

https://apps.apple.com/us/app/waifu2x/id1286485858

I paid extra to get the higher quality model using the in-app purchase option. It crushes the phone's battery life, but runs in only ~10 seconds on an iPhone 13 Pro for a single 1000x1000 input image.

Re: DALL·E now available in beta

#303

I have been having a blast with DALL-E, spending about an hour a day trying out wild combinations and cracking my friends up. I cannot imagine getting bored of it; it's like getting bored with visual stimulus, or art in general. In fact, I've been glad to have a 50/day limit, because it helps me contain my hyperfocus instincts. The information about new pricing is, to me as someone just enjoying making crazy imagines…

MidJourney gives ~unlimited generation for $30/month, and is nearly as good. Unlike DALL-E it doesn't deliberately nerf face generation. I've been having a blast.

Re: DALL·E now available in beta

#304

I have been having a blast with DALL-E, spending about an hour a day trying out wild combinations and cracking my friends up. I cannot imagine getting bored of it; it's like getting bored with visual stimulus, or art in general. In fact, I've been glad to have a 50/day limit, because it helps me contain my hyperfocus instincts. The information about new pricing is, to me as someone just enjoying making crazy imagines…

Sounds kind of like scribblenauts. I would try the craziest things to see what it could come up with.

Re: DALL·E now available in beta

#305

Earlier quoted context omitted.

If accurately reflects the world population then only one in six pictures will be a white person. Half the pictures will be Asian, another sixth will be Indian. Slightly more than half of the pictures will be women. That accurately represents the world's diversity. It won't accurately reflect the world's power balance but that doesn't seem to be their goal. If you want to say "white male CEO" because you want results…

> inoffensive tool. Wouldn't that result end up being like "inoffensive art" or "inoffensive comedy"? Bland, boring and Corporate-PC.

Being offensive is only one way to be interesting.

There are others, like being clever, or being absurd, or being goofy, or being poignant, or being refreshing.

Of the good stuff, offensive humor is only a tiny slice.

Re: DALL·E now available in beta

#306

Wait until someone trains a model like this, for porn. There seems to be a post-DALLE obscenity detector on openAI's tool, as so far I've found it to be entirely robust against deliberate typos designed to avoid simple 'bad word lists'. Ask it for a "pruple violon" and you get purple violins... you get the deal. "Metastable" prompts that may or may not generate obscene (content with nudity, guns, violence as I've fou…

If I had to guess, I'd bet they have a supervised classifier trained to recognize bad content (violence, porn, etc) that they use to filter the generated images before passing them to the user, on top of the bad word lists.

Re: DALL·E now available in beta

#307

Earlier quoted context omitted.

I think people don't realize how huge these models really are. When they're free, it's pretty cool. But charge an amount where there's actual profit in the product? Suddenly seems very expensive and not economically viable for a lot of use cases. We are still in the "you need a supercomputer" phase of these models for now. Something like DALLE mini is much more accessible but the results aren't good enough. Early ear…

What are the resources at work here? What are the resources needed to train this model? If someone just gave you the model for free, what resources would you need to use it to generate new results?

Facebook released over 100 pages of notes a few months ago detailing their training process for a model that is similar in size. Does anyone have a link? I can't seem to find it in my notes, googling links to posts that have been removed or are behind the facebook walled garden.

But I seem to remember they were running 1,000+ 32gb GPUs for 3 months to train it and keeping that infrastructure running day-to-day and tweaking parameters as training continued was the bulk of the 100 pages. It is beyond the reach of anybody but a really big company, at least in the area of very large models, and the large models are where all the recent results are. I wish I was more bullish on algorithm improvements meaning you can get better results on less hardware; there will definitely be some algorithm improvements, but I think we might really need more powerful hardware too. Or pooled resources. Something. These models are huge.

Re: DALL·E now available in beta

#308
post #306

Wait until someone trains a model like this, for porn. There seems to be a post-DALLE obscenity detector on openAI's tool, as so far I've found it to be entirely robust against deliberate typos designed to avoid simple 'bad word lists'. Ask it for a "pruple violon" and you get purple violins... you get the deal. "Metastable" prompts that may or may not generate obscene (content with nudity, guns, violence as I've fou…

If I had to guess, I'd bet they have a supervised classifier trained to recognize bad content (violence, porn, etc) that they use to filter the generated images before passing them to the user, on top of the bad word lists.

Exactly!

Re: DALL·E now available in beta

#309
post #76

I fully expect stock image sites to be swamped by DALL-E generated images that match popular terms (e.g. "business person shaking hands"). Generate the image for $0.15. Sell it for $1.00.

They won't. DALL-E images are mostly not as high quality. The high quality stuff which everyone has been sharing is result of lots of cherry picking.

If the price is low enough, you can have humans rank generated images (maybe using Mechanical Turk or a similar service), and from that ranking choose only the highest quality DALL-E generated images.

Re: DALL·E now available in beta

#310

Earlier quoted context omitted.

What are the resources at work here? What are the resources needed to train this model? If someone just gave you the model for free, what resources would you need to use it to generate new results?

Facebook released over 100 pages of notes a few months ago detailing their training process for a model that is similar in size. Does anyone have a link? I can't seem to find it in my notes, googling links to posts that have been removed or are behind the facebook walled garden. But I seem to remember they were running 1,000+ 32gb GPUs for 3 months to train it and keeping that infrastructure running day-to-day and tw…

> Facebook released over 100 pages of notes a few months ago detailing their training process for a model that is similar in size. Does anyone have a link?

Is https://github.com/facebookresearch/metaseq/blob/main/projec... what you're referring to?

Post reply on HN