Live data from Hacker News

Implementation of Imagen, Google's text-to-image neural network, in PyTorch

github.com

111–120 of 123 posts

Re: Implementation of Imagen, Google's text-to-image neural network, in PyTorch

#111

Earlier quoted context omitted.

You are off by an order of magnitude at least. 256 TPUs-v4 (not pod), would cost you around 20k$/day. They actually used 512 TPUs (256 for base model + 128 for each of the two superresolution models). Assuming an average training time of 1 week as you said, that gives us about 280k$. It's also most likely trained for longer than a week, the base model for Dalle-2 was trained for 100-200k GPU hours, so between 2-4x lo…

Couldn't a bunch of us shell out $5000~$50,000 and do this ourselves? Create a non-profit shell corporation outside US jurisdiction, issue shares, raise funds and open source the result? The shares would simply be votes towards future training dataset endeavors as no profit would be booked here. Say you buy 5000 out of 500,000 shares, that would give you 1% voting power in what dataset to train.

Go right ahead, spend $5-50K of your own money and let us use it for free.

Re: Implementation of Imagen, Google's text-to-image neural network, in PyTorch

#112

Earlier quoted context omitted.

Two reasons: 1) Even though it's all technically very impressive, so far there's not a huge amount of commercialization potential here. OpenAI is charging for its GPT-3 model but its revenue is probably negligible next to the hardware costs (sunk + ongoing) to train it in the first place, let alone the researcher salaries they're paying 2) Most of the stunning examples are cherry-picked. These things fail much more o…

I'm currently working fulltime on AI-powered design suite Accomplice ( https://accomplice.ai ) and if you ask me on a good day I would tell you I do think there's already huge commercial potential. On a bad day, though ;) My current approach is a "model marketplace" ( https://accomplice.ai/models ) where the most popular open source text-to-image models (VQGAN+CLIP, Disco Diffusion, DALL-E Mega coming soon…), sit alo…

*throws money at his screen*

Re: Implementation of Imagen, Google's text-to-image neural network, in PyTorch

#113

Earlier quoted context omitted.

You are off by an order of magnitude at least. 256 TPUs-v4 (not pod), would cost you around 20k$/day. They actually used 512 TPUs (256 for base model + 128 for each of the two superresolution models). Assuming an average training time of 1 week as you said, that gives us about 280k$. It's also most likely trained for longer than a week, the base model for Dalle-2 was trained for 100-200k GPU hours, so between 2-4x lo…

While that is what they did, they also used a batch size of 2048 while training. This is just to speed training up, not a hard requirement. It's easy for Google to justify more money on compute to save engineer iteration loops. I'll have to read the paper for more details, but it would almost certainly cost less (and take longer) to train a model like this in a more resource constrained situation than Google faces .

Increasing batch size does not increase the cost of your training. The opposite actually: With bigger batch size (to an extent), models tend to converge slightly faster so you need less GPU hours. As for the rest, training for 1h with batch size of 2048 on 10 TPUs, or training for 8h with batch size of 256 on 1 TPU has the exact same cost, the cost is just spread over a longer time.

Re: Implementation of Imagen, Google's text-to-image neural network, in PyTorch

#114
post #16

What is the reason Google published their research details about Imagen? Why don't they just keep their findings to themselfes and build products on top of them? Public companies can't do stuff just for the fun of it, right? So there must be some commercial reasoning behind it?

This is an important question.

You're saying it's Google who have done this research. In a way that's true. But really it is Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet and Mohammad Norouzi who did it, with material support from Google.

It's likely that some or all of these people would have refused to do the work they do if Google kept it all as their secret sauce.

And moreover, there are excellent reasons why they wouldn't want to. It's not just the obvious that if it all were secret, they wouldn't be able to use it in their non-Google career advancement. It's also that research without the freedom to talk is far more difficult and frustrating.

On paper, scientific papers are supposed to document the whole of the discovery/innovation. So you might think that an insider, who got to read all the secret Google research papers AND all the public ones would have an advantage. But problem is, even the best written papers with full code and comments inevitably leave out things, especially of the "why this and not that" type.

If you're a researcher in the free world, you can just ask. Especially if you have a public track record of great papers yourself, they will WANT to talk to you. You can learn so much more from the interactive process of back and forth questions than you can from a static piece of information like a scientific paper.

If you work for a secretive and command-driven organization, you need to be careful about what you reveal of your own research when you ask. You can't talk freely. The thought of having to justify your communication to some old-school IBM lawyer type, is going to chill even the most enthusiastic reseacher. It's easier to just stay in your own corporate bubble, and focus on the things your corporation does well since at least you can talk freely to your colleagues (although in really paranoid organizations like the NSA or old IBM, even that may not be true). But then at best you specialize, at worst you fall behind.

Re: Implementation of Imagen, Google's text-to-image neural network, in PyTorch

#115

Earlier quoted context omitted.

In my experience, scraping the data is the easy part. Once you've scraped it you've got to get rid of all the garbage, which is where the issues arise, especially if you're just blindly scraping everything you can find. For example, in a generative model I'm working on, I have a dataset consisting of ~5M images just blindly scraped from a website. After filtering, this drops down to ~500k images, yet a model trained…

Which is why porn is such a great dataset for crowdsource: - lots of people are stimulated by it - lots of people want DALL-E-2 for porn - and lots of people are willing to work towards that common goal The beauty of this is that people are just going to keep coming and coming to it. Like I'm trying to be mature and serious about this. What's it going to take? - Community responsible for scraping dataset, generating…

> people are just going to keep coming and coming to it

ಠ_ಠ

Re: Implementation of Imagen, Google's text-to-image neural network, in PyTorch

#117

How much would it cost to train something like this? Is there even a good dataset for it?

There is a dataset of 5 billion image-text pairs (laion-5b) scraped by various parties. This can then be filtered and used to train these models. Cost is a bit of an issue but there are orgs that have provided compute for open model training. And Imagen is nice because the text encoder part is already available and doesn't need more training, so it would just be the diffusion model components being trained. I'd guess…

Who will do the training? Will the training results be made available publicly? why wont it start until a few weeks from now?

Re: Implementation of Imagen, Google's text-to-image neural network, in PyTorch

#118

Earlier quoted context omitted.

Which jurisdictions are those?

Korea. Hardcore pornography is also illegal there. Keep in mind this is a country where if you leave a bad review after you get scammed by someone with evidence, it is defamation. So not quite leadership the world needs in this industry. Really sad to see ppl on HN flagging all of my comments on this thread. I mean it's not like you can't find celebrity deepfakes including Kpop. The cat is out of the bag and its only…

Also the AI researchers behind these innovations don't want them to be used in this way. And I can't blame them.

Re: Implementation of Imagen, Google's text-to-image neural network, in PyTorch

#119

Earlier quoted context omitted.

I'm currently working fulltime on AI-powered design suite Accomplice ( https://accomplice.ai ) and if you ask me on a good day I would tell you I do think there's already huge commercial potential. On a bad day, though ;) My current approach is a "model marketplace" ( https://accomplice.ai/models ) where the most popular open source text-to-image models (VQGAN+CLIP, Disco Diffusion, DALL-E Mega coming soon…), sit alo…

Your headshots model's outputs look creepy. The eyes are off; e.g. the young african girl. You may want to tune your loss functions to fix that.

Most of the blue eyed ones have the Eyes of Ibad from Dune. (And several of the ones that aren't blue eyed also have weird things going on in the sclerae.)

The two “african american” ones look south or maybe southeast asian (and the one of those that is a “young...girl” looks like a, maybe young, adult.)

All the ones without a racial/ethnic prompt are white, and disproportionately blue eyed (again, including sclerae.)

(It indicate “diverse”, and yet all of the examples read white or Asian, though the unlabeled darker-skinned male figure in the group of six at the top is ambiguous enough to be plausibly be something else.)

The “beautiful woman with curly red hair” has rather radical facial asymmetry, and straight to slightly wavy hair.

Re: Implementation of Imagen, Google's text-to-image neural network, in PyTorch

#120

How much would it cost to train something like this? Is there even a good dataset for it?

There is a dataset of 5 billion image-text pairs (laion-5b) scraped by various parties. This can then be filtered and used to train these models. Cost is a bit of an issue but there are orgs that have provided compute for open model training. And Imagen is nice because the text encoder part is already available and doesn't need more training, so it would just be the diffusion model components being trained. I'd guess…

Are there prefiltered derivatives of Laion-5B available? I can imagine various contraindicated categories you might want to avoid entirely, as well as biases you might want to adjust for by balancing classes in the data (5 billion images gives you a lot of room to balance the dataset).
Post reply on HN