Live data from Hacker News

Imagen: An AI system that creates photorealistic images from input text

imagen.research.google

101–110 of 233 posts

Re: Imagen: An AI system that creates photorealistic images from input text

#101
post #74

Every single time I see an article about a new AI model that has a section called "societal impact" I know immediately they are not releasing the model, the training set, nothing... It seems to be the kind of bullshit statement that those companies put in place of "we paid $500k training this model and we're not giving it for free to anyone".

It's probably a lot more than that, $500k is like one Google ML engineer's salary.

Re: Imagen: An AI system that creates photorealistic images from input text

#102

Earlier quoted context omitted.

The writing issue demonstrably has been solved without any breakthroughs by simply making a larger model (dall-E 2 vs the publicly available dall-e), the same appears to be for other main issues as well - it's just that the publicly available versions are based on smaller/weaker models than the state of art because they're significantly cheaper to run. Also, I seem to recall that at least some models deliberately har…

Neither DALLE version 1 or 2 was completely released. Further, DALLE2 definitely still has issues with generation of text, although latent diffusion can do an okay job of it.

It’s Imagen that can competently generate text, and their ablation studies show it gains the ability at a size somewhat above DALLE2’s, IIRC.

Re: Imagen: An AI system that creates photorealistic images from input text

#103

Is there an AI that creates realistic text from photos?

iOS is pretty good at it if you enable VoiceOver.

Microsoft has an app called Seeing Eye that’s not as good at it.

Both of those use older pre-CLIP technology.

Re: Imagen: An AI system that creates photorealistic images from input text

#104
post #63

I feel like Imagen gets all the noise and people forget about Parti - https://parti.research.google/ It's another google project using a different set of techniques.

That’s because nobody cares that they got a slightly worse result with a different model architecture. Especially the pictures aren’t as sharp, so they’re not as fun to look at.

Parti+Imagen is in development and is competitive again.

Re: Imagen: An AI system that creates photorealistic images from input text

#105

Earlier quoted context omitted.

AFAIK nothing is released yet? They are probably still trying to, like Dall-E 1/2, remove every every image and ban every word that might generate even the slightest hint of controversial imagery before releasing it to the public. To be fair, Stable Diffusion spent months basically doing the same on Discord with thousands of beta users, and had an army of moderators flagging images which were subsequently removed in…

Even if you don’t want to block generating porn, you still want to know if you’re getting it, because nobody wants to get porn when they didn’t ask for it. (and it may be illegal depending on your country) It’s easy to remove the filter from the SD scripts and that’s intentional.

Fair enough and props to Stability for giving people the benefit of the doubt, not to mention actually open sourcing almost all of their work, unlike other "open" projects. Speaking of which, Dall-E bans a decent chunk of the English dictionary in the hopes of preventing people from generating anything even remotely offensive, which can be somewhat frustrating.

Re: Imagen: An AI system that creates photorealistic images from input text

#106

Earlier quoted context omitted.

I kind of get the sentiment about openness but I think it's way more nuanced than you are making out. There are very good reasons for withholding SOTA models, primarily from the info hazard angle and avoiding escalating the capabilities race which is basically the biggest risk we have right now. Google / Deepmind have actually made some good decisions to try and slow down the race (such as waiting to publish).

Capabilities race, seriously? This is not nuclear warfare my guy. It's mathematics.

Information warfare is pretty dangerous too!

Re: Imagen: An AI system that creates photorealistic images from input text

#107
post #57

These tools are amazing for prototyping. I had an idea for a promotional poster, and seeing my idea just by writing it felt like magic. The generated image had too many artifacts to use, but gave me a guideline to follow when creating the real thing in Pixlr. AI content generation (text, image, source code, video, music) will be a huge boon for prototyping where applied judiciously.

Google hasn't released squat . Google's product is vaporware and we shouldn't afford them any airtime until they release something usable. They're just trying to butt in and get press off of the backs of the teams actually working in the open, and that's super lame. Release your model, Google, or stop bragging and talking over the others here. You're greedily sucking oxygen out of the conversation, and as a trillion…

They are afraid of being sued because they are using all the images they have scraped on all the website ever created. They are probably even using images not publicly available.

Re: Imagen: An AI system that creates photorealistic images from input text

#108
post #57

Earlier quoted context omitted.

Google hasn't released squat . Google's product is vaporware and we shouldn't afford them any airtime until they release something usable. They're just trying to butt in and get press off of the backs of the teams actually working in the open, and that's super lame. Release your model, Google, or stop bragging and talking over the others here. You're greedily sucking oxygen out of the conversation, and as a trillion…

I kind of get the sentiment about openness but I think it's way more nuanced than you are making out. There are very good reasons for withholding SOTA models, primarily from the info hazard angle and avoiding escalating the capabilities race which is basically the biggest risk we have right now. Google / Deepmind have actually made some good decisions to try and slow down the race (such as waiting to publish).

They're not slowing down anything. The cat's out of the bag.

What good does a few months lag do when nobody is bracing for impact?

Re: Imagen: An AI system that creates photorealistic images from input text

#109

i only have a passing curiosity in these projects personally. can someone in the field explain why this has exploded recently? there seems to be a lot of these tools released recently (text to image) was there a major breakthrough? a new idea that pushed everyone forward? a recent sharing of talent between groups? edit: just another thought, are they just being posted to HN now, i don't see a date on the page for whe…

Googling the author list, gives me a preprint dated March this year. https://arxiv.org/abs/2205.11487

Re: Imagen: An AI system that creates photorealistic images from input text

#110
post #74

Every single time I see an article about a new AI model that has a section called "societal impact" I know immediately they are not releasing the model, the training set, nothing... It seems to be the kind of bullshit statement that those companies put in place of "we paid $500k training this model and we're not giving it for free to anyone".

I for one am glad that, for once in my life, an obviously huge advancement is taking into account the human impact of releasing the technology responsibly. Maybe it’s just a “BS statement” but given the major strides they’re making at removing racial/gender biases[1] from similar projects, I don’t think it’s just hot air. Especially given the phenomenon of “bias amplification”.

Maybe I’m being too optimistic but either way, given the pace of progress we won’t have to wait very long to play with this magic.

[1] https://openai.com/blog/reducing-bias-and-improving-safety-i...

Post reply on HN