Live data from Hacker News

How Imagen Works

assemblyai.com

91–100 of 104 posts

Re: How Imagen Works

#91
post #69
post #66

Earlier quoted context omitted.

I mean, you try training this thing without a warehouse full of GPUs… to me, the algorithm is just as interesting as the model. Perhaps more so.

"This thing" has already been trained. Nobody is saying the algorithm is not interesting. Just that "this thing" has not been released.

[deleted]

Re: How Imagen Works

#92
post #69
post #66

Earlier quoted context omitted.

I mean, you try training this thing without a warehouse full of GPUs… to me, the algorithm is just as interesting as the model. Perhaps more so.

"This thing" has already been trained. Nobody is saying the algorithm is not interesting. Just that "this thing" has not been released.

Yes, nobody said one way or another. I chose to shine a light on the algorithm. What’s your point?

Re: How Imagen Works

#93
post #73

Earlier quoted context omitted.

>The beautiful thing I'm very much looking forward to be able to play with this tech asap. I'm still excited about AIDungeon. However, OpenAI and creators of the other big name models are restricting the access for good reasons and I'm unsure whether it will be a beautiful thing once it's available for everyone…

Eh, it exists and it's an inevitability that it will eventually be used in terrible ways. OpenAI and Google people just want the CV booster for having created it but want to pretend it's not their fault when it's used to do a racismsexism.

I agree. Tay is still fresh in a lot of folks’ minds.

Re: How Imagen Works

#94
post #20
post #14

Earlier quoted context omitted.

This implementation popped up on hacker news not too long ago. I got it working on Colab first, and then my own GPU at home. But just barely. Need more memory :) https://github.com/lucidrains/imagen-pytorch

Are there any large publicly available models, ready to fine tune and deploy, that were trained on massive data sets? I really want to build services with these.

VQGAN-CLIP

Re: How Imagen Works

#95

I have shown imagen (and dalle2) to a number of people now (non-tech, just everyday friends, family, co-workers) and I have been pretty stunned by the response I get from most people: "Meh, that's kinda cool? I guess?" or "What am I looking at?"..."Ok? So a computer made it? That seems neat" To me I am still trying to get my jaw off the floor from 2 months ago. But the responses have been so muted and shoulder shrugg…

Over a decade ago, Will Wright (of SimCity fame) faked conversational AGI robots in the streets and restaurants of Oakland. It consistently took people 2.4 seconds to go from “Oh look. The robots have arrived.” to “And, I’ll have fries with that.” Hollywood and the media have taught the public that tech is literally magic and can do literally anything. “Anything” is expected and pedestrian.

I often think a similar thing about aliens. That is, instead of the panicking and hysteria or whatever that fiction imagines might accompany the discovery of aliens I fully expect that people will mostly go "Oh, neat. Aliens." And go on with their lives.

Re: How Imagen Works

#96

Earlier quoted context omitted.

I think if you've been paying attentiont to the space, this generation of image diffusion is shocking in how quickly it has improved on what we had a year ago. But if you've never considered that a computer can produce an original image, this is just a new thing computers can do. OTOH I think it's also a lack of imagination in how useful this is, so far the output has been kind of random, so it seems a little gimmick…

You can just type a request in the box if you don't particularly care what the result looks like and also don't care that some of the features might be copyrighted (since large models are quite capable of memorizing their training data.) Asking for two different images in a series that have similar "art styles" is going to be enough work to still need a specialist aka an artist; it'll be most useful in cases you neve…

> Asking for two different images in a series that have similar "art styles" is going to be enough work to still need a specialist aka an artist

Running a separate style transfer network on the generated images is currently possible, although won't achieve the best possible results.

I wouldn't be surprised in the near future to see generation models that can take a text prompt and an image to mimic the style of, which could let it take style into account when generating the image rather than at just the surface level.

Re: How Imagen Works

#97

I have shown imagen (and dalle2) to a number of people now (non-tech, just everyday friends, family, co-workers) and I have been pretty stunned by the response I get from most people: "Meh, that's kinda cool? I guess?" or "What am I looking at?"..."Ok? So a computer made it? That seems neat" To me I am still trying to get my jaw off the floor from 2 months ago. But the responses have been so muted and shoulder shrugg…

I think I can explain this that for most people the whole world is basically magic anyway. They don’t understand any of the details about how any digital tech works so to them they have no framework for which things are impressive and which things are not. The just know that computers can do a great many things that they know nothing about. “Oh I can bank online? Ok.” “Oh, I can have the computer write my book report…

That is exactly what Will Wright (the creator of SimCity and The Sims, and Robot Wars / Battle Bots contestant) was getting at when we made these one-minute robot reality videos about "Empathy" and "Servitude".

His idea was to probe just how much random people on the street (or in a diner) would believe about autonomous intelligent robots operating in the real world.

Of course we were actually hiding behind the scenes tele-operating the robots through hidden cameras and a wireless web interface, listening to what the people said and making the robots respond with a voice synthesizer and sound effects, clicking on pre-written phrases and typing ad-libbed responses.

Empathy (a broken down robot begs for help from passers by on the streets of Oakland):

https://www.youtube.com/watch?v=KXrbqXPnHvE

Servitude (a robot waiter takes orders and serves food in a diner in Oakland, making stupid mistakes and asking for a good review):

https://www.youtube.com/watch?v=NXsUetUzXlg

All his robots aren't as harmless, non-violent, polite, and obsequious as those two. Here's an old interview with Will at Robot Wars 1997:

https://www.youtube.com/watch?v=5nmbs0WqDQM

Here is Super ChiaBot and her MiniBots, created by Will and his daughter Cassidy, getting its leaves shredded and body slammed at BattleBots:

https://www.youtube.com/watch?v=DrArvRG2yQA

Here's a more recent video of Will throwing a tantrum about the failure of SimSandwich, destroying his old creations because they're pixely and poorly rendered, then complaining about how those jerks at EA hate him:

https://www.youtube.com/watch?v=i-7F7s46-9A

Re: How Imagen Works

#98

Earlier quoted context omitted.

I think I can explain this that for most people the whole world is basically magic anyway. They don’t understand any of the details about how any digital tech works so to them they have no framework for which things are impressive and which things are not. The just know that computers can do a great many things that they know nothing about. “Oh I can bank online? Ok.” “Oh, I can have the computer write my book report…

This is the other side of the classic XKCD "Tasks" ( https://xkcd.com/1425/ ). A non-technical person in 2014 (when the above was originally published) would likely have the same conception of the difficulty of recognizing a bird from an image as they would in 2022, even though the task itself has gone from near-insurmountable to off-the-shelf-library in eight years. Even as Imagen and Dall-E 2 amaze us today, these…

The tooltip you get when you hover your cursor over the comic:

"In the 60s, Marvin Minsky assigned a couple of undergrads to spend the summer programming a computer to use a camera to identify objects in a scene. He figured they'd have the problem solved by the end of the summer. Half a century later, we're still working on it."

I'm working with his son Henry Minsky and other great people at Leela AI on that same old problem, applying hybrid symbolic-connectionist constructivist AI by combining neat neural networks with scruffy symbolic logic to understand video, and it's mind boggling what is possible now:

https://leela.ai/

>Our AI system, Leela, is motivated by intrinsic curiosity. Leela creates theories about cause and effect in her world, and then conducts experiments to test these theories. Leela can connect all her knowledge and use this network to make plans, reason about goals, and communicate using grounded natural language.

>Leela has at her core a hybrid symbolic-connectionist network. This means that she uses a dynamic combination of artificial neural networks and symbol networks to learn. Hybrid networks open the door to AI agents that can build their own abstractions on the fly, while still taking full advantage of the power of deep learning.

https://en.wikipedia.org/wiki/Neats_and_scruffies

>Neats and scruffies: Neat and scruffy are two contrasting approaches to artificial intelligence (AI) research. The distinction was made in the 70s and was a subject of discussion until the middle 80s. In the 1990s and 21st century AI research adopted "neat" approaches almost exclusively and these have proven to be the most successful.

>"Neats" use algorithms based on formal paradigms such as logic, mathematical optimization or neural networks. Neat researchers and analysts have expressed the hope that a single formal paradigm can be extended and improved to achieve general intelligence and superintelligence.

>"Scruffies" use any number of different algorithms and methods to achieve intelligent behavior. Scruffy programs may require large amounts of hand coding or knowledge engineering. Scruffies have argued that the general intelligence can only be implemented by solving a large number of essentially unrelated problems, and that there is no magic bullet that will allow programs to develop general intelligence autonomously.

>The neat approach is similar to physics, in that it uses simple mathematical models as its foundation. The scruffy approach is more like biology, where much of the work involves studying and categorizing diverse phenomena.

We're looking for talented engineers and designers to help, including neats and scruffies working together!

https://leela.ai/jobs/

Re: How Imagen Works

#99

Earlier quoted context omitted.

AI achievements will be indistinguishable from human achievements. Humans will try to pass off AI achievements as their own. The line will become so blurred that it will be impossible to tell the difference.

If that happens, all art will simply have no value and art as % of GDP will plummet. Incidentally, this hasn't happened in areas where AI already dominates like chess and go. Magnus Carlsen alone probably generates more "revenue" than all chess AIs combined.

If there was a way to have an AI feed you moves without being caught, then I am positive Carlson wouldn't be at the top for long.

Re: How Imagen Works

#100

Is there a compare and contrast between Imagen and Parti anywhere? I realize the paper came out yesterday, but maybe other people remember what "autoregressive" means better than I do.

Upon first inspection, Parti is not as good. This is perhaps unsurprising - in DALL-E 2 the prior model tested between autoregressive and diffusion models and the diffusion model outperformed
Post reply on HN