Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

621–630 of 661 posts

Re: Imagen, a text-to-image diffusion model

#621
post #550

Earlier quoted context omitted.

I do wonder what Dall-E 2 would output for a request along the lines of "A still life of a vase of flowers in a completely new art style."

Don't have access to Dall-E 2 or Imagen but I do have [1] and [2] locally and they produced [3] with that prompt. [1] https://github.com/nerdyrodent/VQGAN-CLIP.git [2] https://github.com/CompVis/latent-diffusion.git [3] https://imgur.com/a/dCPt35K

Nice. Latent-diffusion has come out very traditional but the VQGAN/CLIP ones are fairly original.

Re: Imagen, a text-to-image diffusion model

#622
post #621

Earlier quoted context omitted.

Don't have access to Dall-E 2 or Imagen but I do have [1] and [2] locally and they produced [3] with that prompt. [1] https://github.com/nerdyrodent/VQGAN-CLIP.git [2] https://github.com/CompVis/latent-diffusion.git [3] https://imgur.com/a/dCPt35K

Nice. Latent-diffusion has come out very traditional but the VQGAN/CLIP ones are fairly original.

From my experiments, the LD one doesn't seem to have been trained on as big or as tagged data set - there's a whole bunch of "in the style of X" that the VQGAN knows* about but the LD doesn't. That might have something to do with it.

Re: Imagen, a text-to-image diffusion model

#623

It’s terrifying that all of these models are one colab notebook away from unleashing unlimited, disastrous imagery on the internet. At least some companies are starting to realize this and are not releasing the source code. However they always manage to write a scientific paper and blog post detailing the exact process to create the model, so it will eventually be recreated by a third party. Meanwhile, Nvidia sees no…

I am absolutely terrified of all this for a different reason: all human professions (not just art) will soon be replaced by “good enough” AI, creating a world flooded with auto-generated junk and billions of people trapped permanently in slums, because you can’t compete with free, and no one can earn a living any longer. It’s an old fear for sure but it seems to be getting closer and closer every day, and yet most of…

As soon as middle class work starts to get automated we will form a new system of resource allocation because all of a sudden the current tax system doesn't work and we go through the mother of all economic crises because no one has any money.

Re: Imagen, a text-to-image diffusion model

#624
post #471

Earlier quoted context omitted.

Just yesterday I was speculating that current AI is bad at math because math on the internet is spectacularly terrible. I think you’re right, and it’s unlikely that we (society) will convince people to label their AI content as such so that scraping is still feasible. It’s far more likely that companies will be formed to provide “pristine training sets of human-created content”, and quite likely they will be subscrip…

>“pristine training sets of human-created content” well, we do have organic/farmed/handcrafted/etc. food. One can imagine information nutrition label - "contains 70% AI generated content, triggers 25% of the daily dopamine release target".

Right? And if you’re a VC backing a new startup you might want to pay that extra to get them started right.

Re: Imagen, a text-to-image diffusion model

#625
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

I think instead the images people want to put on the Internet will do the same for these models as adversarial training did for AlphaZero; it will learn what kinds of images engage human reaction.

Re: Imagen, a text-to-image diffusion model

#626

Earlier quoted context omitted.

For comparison, most humans can't draw a bicycle: https://www.wired.com/2016/04/can-draw-bikes-memory-definite...

I blame it on the surprisingly structural cleverness of a bicycle. Opposing triangles probably isn’t the first thing most people think of when they think of a bicycle (vs two wheels and some handlebars)

They also can't draw pennies, the letter 'g' with the loop, and so on (https://www.gwern.net/docs/psychology/illusion-of-depth/inde...). Bicycles may be clever, but the shallowness of mental representation is real.

Re: Imagen, a text-to-image diffusion model

#627

Earlier quoted context omitted.

That's exactly what's happening. Doing the search from the article of "unprofessional hair for work" brings up images with headlines like "It's ridiculous to say that black women's hair is unprofessional". (In addition to now bringing up images from that article itself and other similar articles comparing Google Images searches.)

You’re getting cause and effect backwards. The coverage of this changed the results, as did Google’s ensuing interventions.

I don't think so. You can set the search options to only find images published before the article, and even find some of the original images.

One image links to the 2015 article, "It's Ridiculous To Say Black Women's Natural Hair Is 'Unprofessional'!". The Guardian article on the Google results is from 2016.

Another image has the headline, "5 Reasons Natural Hair Should NOT be Viewed as Unprofessional - BGLH Marketplace" (2012).

Another: "What to Say When Someone Calls Your Hair Unprofessional".

Also, have you noticed how good and professional the black women in the Guardian's image search look? Most of them look like models with photos taken by professional photographers. Their hair is meticulously groomed and styled. This is not the type of photo an article would use to show "unprofessional hair". But it is the type of photo the above articles opted for.

Re: Imagen, a text-to-image diffusion model

#628
post #581

Earlier quoted context omitted.

No

To expand a bit for the grandparent, if you check out this authors other repos you'll notice they have a thing for implementing these papers (multiple DALLE-2 implementations for instance). You should expect to see an implementation there pretty quickly I'd guess.

Thank you. I just saw a GitHub repo, empty except for a citation and a claim that it was an implementation of Imagen, and thought it was perhaps some satirical statement about open source or something. With the context it makes a lot more sense.

Re: Imagen, a text-to-image diffusion model

#629

It’s terrifying that all of these models are one colab notebook away from unleashing unlimited, disastrous imagery on the internet. At least some companies are starting to realize this and are not releasing the source code. However they always manage to write a scientific paper and blog post detailing the exact process to create the model, so it will eventually be recreated by a third party. Meanwhile, Nvidia sees no…

I am absolutely terrified of all this for a different reason: all human professions (not just art) will soon be replaced by “good enough” AI, creating a world flooded with auto-generated junk and billions of people trapped permanently in slums, because you can’t compete with free, and no one can earn a living any longer. It’s an old fear for sure but it seems to be getting closer and closer every day, and yet most of…

[deleted]

Re: Imagen, a text-to-image diffusion model

#630
post #463

Earlier quoted context omitted.

The irony is that when the majority of content becomes computer-generated, most of that content will also be computer-consumed. Neil Stephenson covered this briefly in "Fall; or Dodge In Hell." So much 'net content was garbage, AI-generated, and/or spam that it could only be consumed via "editors" (either AI or AI+human, depending on your income level) that separated the interesting sliver of content from...everythin…

He was definitely onto something in that book where people also resort to using blockchains to fingerprint their behavior and build an unbreakable chain of authenticity. Later in that book that is used to authorize the hardware access of the deceased and uploaded individuals. A bit far out there in terms of plot but the notion of authenticating based on a multitude of factors and fingerprints is not that strange. We'…

As usual, Stephenson is at his best when he's taking current trends and extrapolating them to almost absurd extremes...until about a decade passes and you realize they weren't that extreme after all.

I loved that he extended the concept of identity as an individualized pattern of events and activities to the real world: the innovation of face masks with seemingly random but unique patterns to foil facial recognition systems but still create a unique identity.

Like you say, the story itself had horrible flaws (I'm still not sure if I liked it in its totality, and I'm a Stephenson fan since reading Snow Crash on release in '92), but still had fascinating and thought provoking content.

Post reply on HN