Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

461–470 of 661 posts

Re: Imagen, a text-to-image diffusion model

#461
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

Huh, I had never thought of that. Makes it seem like there's a small window of authenticity closing. The irony is that if you had a great discriminator to separate the wheat from the chaff, that it would probably make its way into the next model and would no longer be useful. My only recommendation is that OpenAI et al should be tagging metadata for all generated images as synthetic. That would be a really interestin…

The OpenAI access agreement actually says that you must add (or keep?) a watermark on any generated images, so you’re in good company with that line of thinking.

Re: Imagen, a text-to-image diffusion model

#462

I know that some monstrous majority of cognitive processing is visual, hence the attention these visually creative models are rightfully getting, but personally I am much more interested in auditory information and would love to see a promptable model for music. Was just listening to "Land Down Under" from Men At Work. Would love to be able to prompt for another artist I have liked: "Tricky playing Land Down Under."…

I agree. How cool would it be to get an 8 min version of your favorite song? Or an instant DnB remix? Or 10 more songs in the style of your favorite album?

You can sort of do that with https://fairuseify.ml

Re: Imagen, a text-to-image diffusion model

#463
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

The irony is that when the majority of content becomes computer-generated, most of that content will also be computer-consumed.

Neil Stephenson covered this briefly in "Fall; or Dodge In Hell." So much 'net content was garbage, AI-generated, and/or spam that it could only be consumed via "editors" (either AI or AI+human, depending on your income level) that separated the interesting sliver of content from...everything else.

Re: Imagen, a text-to-image diffusion model

#464
post #402

Earlier quoted context omitted.

I believe we’re lacking someone training up a large music model here, but GPT-style transformers can produce music. gwern can maybe comment here. An actually scary thing is that AIs are getting okay at reproducing people’s voices.

Voice synthesis has been going steady. Lots of commercial and hobbyist interest: you can use 15.ai for crackerjack SaaS voice synthesis in a slick free UI; and if you want to run the models yourselves, Tortoise just released a FLOSS stack of remarkable quality. Music, I'm afraid, appears stuck in the doldrums of small one-offs doing stuff like MIDI. Nothing like the breadth & quality of Jukebox has come out since it,…

The developer behind Tortoise is experimenting with using diffusion for music generation:

https://nonint.com/2022/05/04/friends-dont-let-friends-train...

Re: Imagen, a text-to-image diffusion model

#465
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

Look at carpentry blogs, recipe blogs. Nearly all of it is junk content. I bet if you combined GPT and imagen or dalle2 you could replace all of them. Just provide a betty crocker recipe and let it generate a blog that has weekly updates and even a bunch of images - "happy family enjoying pancakes together" I can see the future as being devoid of any humanity.

I wrote a comedic "Best Apache Chef recipe" article[1] mocking these sites.

I guess the concern would be: If one of these recipe websites _was_ generated by an AI, the ingredients _look_ correct to an AI but are otherwise wrong - then what do you do? Baking soda swapped with baking powder. Tablespoons instead of teaspoons. Add 2tbsp of flower to the caramel macchiato. Whoops! Meant sugar.

[0] http://slimsag.com/best-apache-chef-recipe/1438731.htm

Re: Imagen, a text-to-image diffusion model

#466
post #250

Earlier quoted context omitted.

That only lasts until the community copies the paper and catches up. For example the open source DALLE-2 implementation is coming along great: https://github.com/lucidrains/DALLE2-pytorch

Imagen actually shows some of the components in DALLE2 is unnecessary, so Imagen will end up being easier to build. I'll definitely add the dynamic thresholding trick from Imagen to DALLE2 repository though; that is a finding that should boost any DDPMs using classifier free guidance.

Thanks for all your work on these projects!

Re: Imagen, a text-to-image diffusion model

#467
post #384

Earlier quoted context omitted.

I firmly believe that ~20-40% of the machine learning community will say that all ML models are dumb statistical interpolators all the way until a few years after we achieve AGI. Roughly the same groups will also claim that human intelligence is special magic that cannot be recreated using current technology. I think it’s in everyone’s benefit if we start planning for a world where a significant portion of the expert…

> The dangers of dismissing the possibility of AGI emerging in the next 5-10 years are huge. Again, I think we should consider "The Human Alignment Problem" more in this context. The transformers in question are large, heavy and not really prone to "recursive self-improvement". If the ML-AGI works out in a few years, who gets to enter the prompts?

A DAO.

Re: Imagen, a text-to-image diffusion model

#468

Earlier quoted context omitted.

https://www.vice.com/en/article/93ywpp/text-adventure-game-c... TL;DR generative story site creators employ human moderation after horny people inevitably use site to make gross porn; horny people using site to make regular porn justifiably freaked out Bring your popcorn

AI is for porn

Time to update this song? https://www.youtube.com/watch?v=j6eFNRKEROw

Re: Imagen, a text-to-image diffusion model

#469

Earlier quoted context omitted.

What is the AI dungeon incident?

https://www.vice.com/en/article/93ywpp/text-adventure-game-c... TL;DR generative story site creators employ human moderation after horny people inevitably use site to make gross porn; horny people using site to make regular porn justifiably freaked out Bring your popcorn

I always felt like the AI Dungeon response is absurdly stupid. Yeah, erotic text involving minors is not exactly something you want to be associated with, but I've heard that one of their approaches to avoiding it was to ban numbers below 18. That's... not worth it.

I feel like it would've been more than reasonable for them to have taken the position that the AI might output something distasteful, and implement a filter for people who were afraid of it.

Re: Imagen, a text-to-image diffusion model

#470

Metacalculus, a mass forecasting site, has steadily brought forward the prediction date for a weakly general AI. Jaw-dropping advances like this, only increase my confidence in this prediction. "The future is now, old man." https://www.metaculus.com/questions/3479/date-weakly-general...

"The future is already here — It’s just not very evenly distributed"
Post reply on HN