Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

451–460 of 661 posts

Re: Imagen, a text-to-image diffusion model

#451
post #132

Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.

Ethics,racist, LGBT, bla,bla. If we talk about political correct, I really suggest you guys go somewhere else. instead of stay in hacker news. AI generated porn is better that let some people, who do not want to do the porn, doing the porn themselves.

Re: Imagen, a text-to-image diffusion model

#452
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

Look at carpentry blogs, recipe blogs. Nearly all of it is junk content. I bet if you combined GPT and imagen or dalle2 you could replace all of them. Just provide a betty crocker recipe and let it generate a blog that has weekly updates and even a bunch of images - "happy family enjoying pancakes together" I can see the future as being devoid of any humanity.

The future digital landscape might be void of humanity, but there will still be real humans living next door to you ;)

Re: Imagen, a text-to-image diffusion model

#453

Earlier quoted context omitted.

Yep, that's the hard problem Google is not comfortable releasing the API to this until they have it solved.

But why is it a problem? The AI is just a mirror showing us ourselves. That’s a good thing. How does it help anyone to make an AI that presents a fake world so that we can pretend that we live in a world that we actually don’t? Disassociation from reality is more dangerous than bias.

The AI is a mirror of the text and image corpora it was presented, as parsed and sanitized by the team in question.

Re: Imagen, a text-to-image diffusion model

#454

I thought I was doing well after not being overly surprised by DALL-E 2 or Gato. How am I still not calibrated on this stuff? I know I am meant to be the one who constantly argues that language models already have sophisticated semantic understanding, and that you don't need visual senses to learn grounded world knowledge of this sort, but come on, you don't get to just throw T5 in a multimodal model as-is and have i…

I haven't been overly surprised by any of it. The final product is still the same, no matter how much they scale it up.

All of these models seem to require a human to evaluate and edit the results. Even Co-Pilot. In theory this will reduce the number of human hours required to write text or create images. But I haven't seen anyone doing that successfully at scale or solving the associated problems yet.

I'm pessimistic about the current state of AI research. It seems like it's been more of the same for many years now.

Re: Imagen, a text-to-image diffusion model

#455

Interesting to me that this one can draw legible text. DALLE models seem to generate weird glyphs that only look like text. The examples they show here have perfectly legible characters and correct spelling. The difference between this and DALLE makes me suspicious / curious. I wish I could play with this model.

Still has the issue with screwing up mechanical objects. In their demo checkout the wheels on the skateboards, all over the place.

For comparison, most humans can't draw a bicycle:

https://www.wired.com/2016/04/can-draw-bikes-memory-definite...

Re: Imagen, a text-to-image diffusion model

#456
post #384

Earlier quoted context omitted.

I firmly believe that ~20-40% of the machine learning community will say that all ML models are dumb statistical interpolators all the way until a few years after we achieve AGI. Roughly the same groups will also claim that human intelligence is special magic that cannot be recreated using current technology. I think it’s in everyone’s benefit if we start planning for a world where a significant portion of the expert…

> The dangers of dismissing the possibility of AGI emerging in the next 5-10 years are huge. Again, I think we should consider "The Human Alignment Problem" more in this context. The transformers in question are large, heavy and not really prone to "recursive self-improvement". If the ML-AGI works out in a few years, who gets to enter the prompts?

Me.

... ... ...

Obviously "/s", obviously joking, but meant to highlight that there are a few parties that would all answer "me" and truly mean it, often not in a positive way.

Re: Imagen, a text-to-image diffusion model

#457
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

Just yesterday I was speculating that current AI is bad at math because math on the internet is spectacularly terrible.

I think you’re right, and it’s unlikely that we (society) will convince people to label their AI content as such so that scraping is still feasible.

It’s far more likely that companies will be formed to provide “pristine training sets of human-created content”, and quite likely they will be subscription based.

Re: Imagen, a text-to-image diffusion model

#458

Earlier quoted context omitted.

You know, it wouldn't surprise me if people talking about how black curly hair shouldn't be seen as unprofessional contributed to google thinking there's an association between the concepts of "unprofessional hair" and "black curly hair"

You really are not helping that cause. As a foreigner[], your point confused me anyway, and doing a Google for cultural stuff usually gets variable results. But I did laugh at many of the comments here https://www.reddit.com/r/TooAfraidToAsk/comments/ufy2k4/why_... [] probably, New Zealand, although foreigner is relative

Haha. I've got some personal experience with that one. I used to live in a house with many other people, and one girl was rastafarian and from jamacia and had dreadlocks, and another girl in the house (who wasn't black) thought that her hairstyle was very offensive. We had to have several conflict resolution meetings about it.

As silly as it seemed, I do think everyone is entitled to their own opinion and I respect the anti-dreadlocks girl for standing up for what she believed in even when most people were against her.

Re: Imagen, a text-to-image diffusion model

#459
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

How would that really happen? It seems to me you're assuming that there's no such thing as extant databases of actual oil paintings, that people will stop producing, documenting, and curating said paintings. I think the internet and curated image databases are far more well kept than your proposed model accounts for.

Re: Imagen, a text-to-image diffusion model

#460
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

Look at carpentry blogs, recipe blogs. Nearly all of it is junk content. I bet if you combined GPT and imagen or dalle2 you could replace all of them. Just provide a betty crocker recipe and let it generate a blog that has weekly updates and even a bunch of images - "happy family enjoying pancakes together" I can see the future as being devoid of any humanity.

I’d much rather skip the blog format and replace them with an AI that can answer “Please provide a pie recipe like my grandparent’s”, or “I’d like to make these ribs on the BBQ so that they come out flavourful, soft, and a little sweet.”
Post reply on HN