Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

511–520 of 661 posts

Re: Imagen, a text-to-image diffusion model

#511
post #473

One thing that no one predicted in AI development was how good it would become at some completely unexpected tasks while being not so great at the ones we supposed/hoped it would be good. AI was expected to grow like a child. Somehow blurting out things that would show some increasing understanding on a deep level but poor syntax. In fact we get the exact opposite. AI is creating texts that are syntaxically correct a…

Syntactically* I know it's the most trivial of things, but in case you were curious as I often am!

Re: Imagen, a text-to-image diffusion model

#512
post #473

One thing that no one predicted in AI development was how good it would become at some completely unexpected tasks while being not so great at the ones we supposed/hoped it would be good. AI was expected to grow like a child. Somehow blurting out things that would show some increasing understanding on a deep level but poor syntax. In fact we get the exact opposite. AI is creating texts that are syntaxically correct a…

Syntactically* I know it's the most trivial of things, but in case you were curious as I often am!

Oh God! Am I a bot?

Re: Imagen, a text-to-image diffusion model

#514
post #132

Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.

check out open source alternative dalle-mini: https://huggingface.co/spaces/dalle-mini/dalle-mini

Thanks for the link. Doesn't appear to be very good yet, though.

Re: Imagen, a text-to-image diffusion model

#515
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

Eventually the only jobs humans will have is training AI to act human. Sounds very Philip K Dick now that I think about it.

Re: Imagen, a text-to-image diffusion model

#516
post #493
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

People training newer models just have to look for the "Imagen" tag or the Dall-E2 rainbow at the corner and heuristically exclude images having these. This is trivial. Unless you assume there are bad actors who will crop out the tags. Not many people now have access to Dall-E2 or will have access to Imagen. As someone working in Vision, I am also thinking about whether to include such images deliberately. Using imag…

In my melody generation system I'm already including melodies that I've judged as "good" (https://www.youtube.com/playlist?list=PLoCzMRqh5SkFwkumE578Y...) in the updated training set. Since the number of catchy melodies that have been created by humans is much, much lower than the number of pretty images, it makes a significant difference. But I'd expect that including AI-generated images without human quality judgement scores in the training set won't be any better than other augmentation techniques.

Re: Imagen, a text-to-image diffusion model

#517

Earlier quoted context omitted.

Still has the issue with screwing up mechanical objects. In their demo checkout the wheels on the skateboards, all over the place.

For comparison, most humans can't draw a bicycle: https://www.wired.com/2016/04/can-draw-bikes-memory-definite...

I blame it on the surprisingly structural cleverness of a bicycle. Opposing triangles probably isn’t the first thing most people think of when they think of a bicycle (vs two wheels and some handlebars)

Re: Imagen, a text-to-image diffusion model

#518

Would be fascinated to see the DALL-E output for the same prompts as the ones used in this paper. If you've got DALL-E access and can try a few, please put links as replies!

Posting a few comparisons here. https://twitter.com/joeyliaw/status/1528856081476116480?s=21...

Looking at these… I can’t help but wonder if these are literal examples of AI imagination?

Re: Imagen, a text-to-image diffusion model

#519
post #503

Earlier quoted context omitted.

Posting a few comparisons here. https://twitter.com/joeyliaw/status/1528856081476116480?s=21...

Imagen seems more realistic where Dall-E2 is more feel-good . That is what I feel personally.

I agree with you, but for me, Dall·E 2 feels good because 90% of the time I can keep hitting the generate button and massage the prompt until I get something inspirational, surprisingly, or visually pleasing. Without access to Imagen, it's impossible for me to compare how much of the "realistic feels" of its images is constrained by the taste of the cherry-pickers.

Re: Imagen, a text-to-image diffusion model

#520

Earlier quoted context omitted.

Posting a few comparisons here. https://twitter.com/joeyliaw/status/1528856081476116480?s=21...

Looking at these… I can’t help but wonder if these are literal examples of AI imagination?

I've started to ask myself if my own creativity is a result of random sampling from the diffusion tapestry of associated memories and experience on that topic.
Post reply on HN