I apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself. Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are. Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication . And short of FAIR? D2 lacrosse. There are exceptions to such…
Imagen, a text-to-image diffusion model
371–380 of 661 posts
Re: Imagen, a text-to-image diffusion model
#372I thought I was doing well after not being overly surprised by DALL-E 2 or Gato. How am I still not calibrated on this stuff? I know I am meant to be the one who constantly argues that language models already have sophisticated semantic understanding, and that you don't need visual senses to learn grounded world knowledge of this sort, but come on, you don't get to just throw T5 in a multimodal model as-is and have i…
They’re all fundamentally anthropocentric: people argue until they are blue in the face about what “intelligent” means but it’s always implicit that what they really mean is “how much like me is this other thing”.
Language models, even more so than the vision models that got them funded have empirically demonstrated that knowing the probability of two things being adjacent in some latent space is at the boundary indistinguishable from creating and understanding language.
I think the burden is on the bright hominids with both a reflexive language model and a sex drive to explain their pre-Copernican, unique place in the theory of computation rather than vice versa.
A lot of these problems just aren’t problems anymore if performance on tasks supersedes “consciousness” as the thing we’re studying.
Re: Imagen, a text-to-image diffusion model
#373Earlier quoted context omitted.
> I mean, from my perspective, the skill in these (and DALL-E's) image reproductions is truly astonishing. A basic part of it is that neural networks combine learning and memorizing fluidly inside them, and these networks are really really big, so they can memorize stuff good. So when you see it reproduce a Shiba Inu well, don’t think of it as “the model understands Shiba Inus”. Think of it as making a collage out of…
> these networks are really really big, so they can memorize stuff good. People tend to really underestimate just how big these models are. Of course these models aren't simply "really really big" MLPs, but the cleverness of the techniques used to build them is only useful at insanely large scale. I do find these models impressive as examples of "here's what the limit of insane amounts of data, insane amounts of comp…
I think it's fair to say that this is the way it's always been. In 1990, you couldn't hack on an accurate fluid simulation at home, you needed to be at a university or research lab with access to a big cluster. But then, 10 years later, you could do it on a home PC. And then, 10 years after that, you could do it in a browser on the internet.
It's the same with this AI stuff.
I think if we weren't in the midst of this unique GPU supply crunch, the price of a used 1070 would be about $100 right now -- such a card would be state of the art 10 years ago!
Re: Imagen, a text-to-image diffusion model
#374Interesting to me that this one can draw legible text. DALLE models seem to generate weird glyphs that only look like text. The examples they show here have perfectly legible characters and correct spelling. The difference between this and DALLE makes me suspicious / curious. I wish I could play with this model.
Re: Imagen, a text-to-image diffusion model
#375Their slider with examples at the top showed a prompt along the lines of "a chrome plated duck with a golden beak confronting a turtle in a forest" and the resulting image was perfect - except the turtle had a golden shell.
Re: Imagen, a text-to-image diffusion model
#376Earlier quoted context omitted.
The bulk of the trained data is from western technology, images, books, television, movies, photography, media. That's where the very real and recognized biases come from. They're the result of a gap in data nothing more. Look at how DALL-E 2 produces little bears rather than bear sized bears. Because its data doesn't have a lot of context for how large bears are. So you wind up having to say "very large bear" to DAL…
That's true for some things, but the "gender bias for some professions" is likely to just be reflecting reality.
Re: Imagen, a text-to-image diffusion model
#377Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.
Re: Imagen, a text-to-image diffusion model
#378Earlier quoted context omitted.
Not the person you responded to, but I do see how someone could be hurt by that, and I want to avoid hurting people. But is this the level at which we should do it? Could skewing search results, i.e. hiding the bias of the real world, give us the impression that everything is fine and we don't need to do anything to actually help people? I have a feeling that we need to be real with ourselves and solve problems and n…
>Could skewing search results, i.e. hiding the bias of the real world Your logic seems to rest on this assumption which I don't think is justified. "Skewing search results" is not the same as "hiding the biases of the real world". Showing the most statistically likely result is not the same as showing the world how it truly is. A generic nurse is statistically going to be female most of the time. However, a model tha…
Every model will have some random biases. Some of those random biases will undesirably exaggerate the real world. Every model will undesirably exaggerate something. Therefore no model should be shared.
Your goal is nice, but impractical?
Re: Imagen, a text-to-image diffusion model
#379https://github.com/lucidrains/imagen-pytorch
Re: Imagen, a text-to-image diffusion model
#380Earlier quoted context omitted.
Not elitist at all; I highly appreciate this post. I know the basics of ML but otherwise am clueless when it comes to the true depths of this field and it's interesting to hear this perspective.
I used a lot of jargon and lingo and inside baseball in that post, it was intended for people who have deep background. But if you’re interested I’m happy to (attempt) answers to anything that was jargon: by virtue of HN my answers will be peer-reviewed in real time, and with only modest luck, a true expert might chime in.