Live data from Hacker News

DALL·E now available in beta

openai.com

421–430 of 579 posts

Re: DALL·E now available in beta

#421
post #182

Earlier quoted context omitted.

DALLE images are still only 1024 px wide. Which has its uses, but I don’t think the stock photo industry is in real danger until someone figures out a better AI superresolution system that can produce larger and more detailed images.

DallE2 + Topaz Gigapixel AI works amazingly well.

Wow

https://www.topazlabs.com/gigapixel-ai

No kidding.

Re: DALL·E now available in beta

#422
post #28

Something I haven’t seen anyone talking about with these huge models: how do future models get trained when more content online is model generated to start with? Presumably you don’t wanna train a model on autogenerated images or text, but you can’t necessarily know which is which.

This precise thing is causing a funny problem in specialty areas. People are using e.g. Google Lens to identify plants, birds and insects, which sometimes returns wrong answers e.g. say it sees a picture of a Summer Tanager and calls it a Cardinal. If the people then post "Saw this Cardinal" and the model picks up that picture/post and incorporates it into its training set, it's just reinforcing the wrong identificat…

Imagine when there is an AI that is monitoring content creation and keeping tabs of original sources....

Re: DALL·E now available in beta

#423

Since many people will start generating their first images soon, be sure to check out this amazing DALL-E prompt engineering book [0]. It will help you get the most out of DALL-E. [0]: https://dallery.gallery/wp-content/uploads/2022/07/The-DALL%... (PDF)

thanks! love the link highlight

Re: DALL·E now available in beta

#424

How do you interface with DALL-E? For MidJourney I was painfully surprised to find that everything is done through chat messages on a Discord server. I'm not a paid member, so I have to enter my prompts in public channels. It's extremely easy to lose your own prompts in the rapidly flowing stream of prompts going on. I can kind of see why they did it that way--maybe, if I squint really hard--to try to promote visibil…

Yeah that's definitely one of the worst aspects of using midjourney, supposedly a API is coming but it doesn't look like it's going to be happening anytime soon.

I don't know who thought that discord would make a good GUI front end...

Re: DALL·E now available in beta

#425
post #306

Wait until someone trains a model like this, for porn. There seems to be a post-DALLE obscenity detector on openAI's tool, as so far I've found it to be entirely robust against deliberate typos designed to avoid simple 'bad word lists'. Ask it for a "pruple violon" and you get purple violins... you get the deal. "Metastable" prompts that may or may not generate obscene (content with nudity, guns, violence as I've fou…

If I had to guess, I'd bet they have a supervised classifier trained to recognize bad content (violence, porn, etc) that they use to filter the generated images before passing them to the user, on top of the bad word lists.

This is mentioned, "content filters" are "blocking images that violate our content policy — which does not allow users to generate violent, adult, or political content, among other categories" and they "limited DALL·E’s exposure to these concepts by removing the most explicit content from its training data."

Re: DALL·E now available in beta

#426
post #316

Earlier quoted context omitted.

My heart sank when I saw the pricing model. I’ve been creating generative art since 2016 and I’ve been anxiously waiting for my invite. I wont be able to afford to generate the volume of images it takes to get good ones at this price point. I can afford $20/mo for something like this but I just can’t swing $200 to $300 it realistically takes to get interesting art out of these CLIP-centric models. Heck, the initial 5…

If you’re technically inclined, I urge you to explore some newer Colabs being shared in this space. They offer vastly more configurable tools, work great for free on Google Colab, are straightforward to run on a local machine. Meanwhile we should prepare ourselves for a future where the best generative models cost a lot more as these companies slice and dice the (huge) burgeoning market here.

Can you share a few, please? I am already out of credits on Midjourney and I don't even feel like I got the hang of it until I was almost out.

Re: DALL·E now available in beta

#427

The name "OpenAI" to me implies being open-source. I have an RTX 3080 and will likely be buying a 4090 when it comes out. Will I ever be able to generate these images locally, rather than having to use a paid service? I've done it with DALL-E Mini, but the images from that don't hold a candle to what DALL-E 2 produces.

You should show up to the US Open with a tennis racket next year and see if they'll let you have a go, too.

Are you an AI as well? Because within the context of tech "Open" definitely has that connotation

Re: DALL·E now available in beta

#428
post #413

Earlier quoted context omitted.

> I think people don't realize how huge these models really are. They really aren't that large by the contemporary scaling race standards. DALLE-2 has 3.5B parameters, which should fit on an old GPU like Nvidia RTX2080, especially if you optimize your model for inference [1][2] which is commonly done by ML engineers to minimize costs. With optimized model, your memory footprint is ~1 byte per parameter, and some less…

It's true that image models are much less of a burden on GPU VRAM than a model like BLOOM where fitting it into a few A100s is ideal, but these diffusion models are a PITA for a ordinary hobbyist in terms of total compute: the CLIP pass over the text input is almost free, but then you feed it into the diffusion model, for one sample you'll be doing 10-100 forward passes (depending on how fancy the diffusion methods a…

All valid points, of course. As an independent explorer I adapted my workflow to use night's worth of workstation compute to generate a crop of new images from a simple templated prompt language. It also helps a lot to have at least two presets - "exploratory" and "hq", to minimize iteration time and maximize quality of promising prompts.

Still, I think optimization of diffusion models for efficient inference isn't yet pushed to the limits. At least if we look at what's available to the public - AFAIK public inference software distributions didn't even quantize their weights.

Re: DALL·E now available in beta

#429
post #58

I was supposed to be making a video game, but got a bit sidetracked when DALL·E came out and made this website on the side: http://dailywrong.com/ (yes I should get SSL). It's like The Onion, but all the articles are made with GPT-3 and DALL·E. I start with an interesting DALL·E image, then describe it to GPT-3 and ask it for an Onion-like article on the topic. The results are surprisingly good.

I’m curious, if they’re only making DALL-E accessible now, and if GPT-3 was never really accessible (as far as I know). How do you have access to these things to generate text and images?

Re: DALL·E now available in beta

#430

Surprised by the lack of comments on the ethics of DALL-E being trained on artists content whereas copilot threads are chock full of devs up in arms over models trained on open source code. Isn’t it the same thing?

I recently talked with a concept artist about DALL-E and first thing they mentioned was "you know that's all stolen art, right?" Immediately made me think of GitHub Copilot.

However the artists being featured in DALL-E's newsletters can't stop gushing about 'the new instrument they are learning how to play' and other such metaphors that are meant to launder what's going on.

My theory is that the professions most at-risk for automation are acting on their anxieties. Must not be a lot of freelance artists on HN, and a whole lot of programmers.

I think the artists have an even clearer case. I don't think GitHub Copilot is ready to steal anyone's job yet. But DALL-E is poised to replace all formerly commissioned filler art for magazines, marketing sites, and blogs. Now the only point to hiring a human is to say you hired a human. Our filler art is farm-to-table.

Post reply on HN