Earlier quoted context omitted.
How can we prepare for this? This will result in mass social unrest.
I think the serious answer is that it is yet another labor multiplier like electricity and software. Our tech since the industrial revolution has allowed us to elevate ourselves from a largely agrarian society to space and cyberspace. AI, by all appearances, continues to be a tool, just the latest in a long line of better tools. It still requires a human to provide intent and direction. Right now in my job, I command…
Imagen, a text-to-image diffusion model
391–400 of 661 posts
Re: Imagen, a text-to-image diffusion model
#392Earlier quoted context omitted.
> I mean, from my perspective, the skill in these (and DALL-E's) image reproductions is truly astonishing. A basic part of it is that neural networks combine learning and memorizing fluidly inside them, and these networks are really really big, so they can memorize stuff good. So when you see it reproduce a Shiba Inu well, don’t think of it as “the model understands Shiba Inus”. Think of it as making a collage out of…
> these networks are really really big, so they can memorize stuff good. People tend to really underestimate just how big these models are. Of course these models aren't simply "really really big" MLPs, but the cleverness of the techniques used to build them is only useful at insanely large scale. I do find these models impressive as examples of "here's what the limit of insane amounts of data, insane amounts of comp…
Other funding models are possible as well, in the grand scheme of things the price for these models is small enough.
Re: Imagen, a text-to-image diffusion model
#393Probably just a frontend coding mistake, and not an error in the model, but in the interactive example if you select: "A photo of a Shiba Inu dog Wearing a (sic) sunglasses And black leather jacket Playing guitar In a garden" The Shiba Inu is not playing a guitar.
Re: Imagen, a text-to-image diffusion model
#394Interesting discovery they made > We show that scaling the pretrained text encoder size is more important than scaling the diffusion model size. There seems to be an unexpected level of synergy between text and vision models. Can't wait to see what video and audio modalities will add to the mix.
Particularly as you approach the point where the image quality itself is superb and people increasingly turn to attacking the semantics & control of the prompt to degrade the quality ("...The donkey is holding a rope on one end, the octopus is holding onto the other. The donkey holds the rope in its mouth. A cat is jumping over the rope..."). For that sort of thing, it's hard to see how simply beefing up the raw pixel-generating part will help much: if the input seed is incorrect and doesn't correctly encode a thumbnail sketch of how all these animals ought to be engaging in outdoors sports, there's nothing some low-level pixel-munging neurons can do to help much.
Re: Imagen, a text-to-image diffusion model
#395Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.
Granted that's a selection bias: you likely won't hear about the cases where legit obscene output occurs. (the only notable case I've heard is the AI Dungeon incident)
Re: Imagen, a text-to-image diffusion model
#396Earlier quoted context omitted.
> Copenhagen ethics (used by most people) The idea that most people use any coherent ethical framework (even something as high level and nearly content-free as Copenhagen) much less a particular coherent ethical framework is, well, not well supported by the evidence. > require that all negative outcomes of a thing X become yours if you interact with X. It is not sensible to interact with high negativity things unless…
I'm sure you are capable of steelmanning the argument.
“There exists an ethical framework—not the Copenhagen interpretation —to which some minority of the population adheres in which trying and failing to a correct a problem incurs retroactive blame for the existence of the problem but seeing it and just saying ‘sucks, but not my problem’ does not,“ is probably true, but not very relevant.
It's logical for Google to avoid involvement with porn, and to be seen doing so, because even though porn is popular involvement with it is nevertheless politically unpopular, and Google’s business interest is in not making itself more attractive as a political punching bag. The popularity of Copenhagen ethics (or their distorted cousins) don't really play into it, just self interest.
Re: Imagen, a text-to-image diffusion model
#397Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.
Running inference on one of these models takes like a GPU minute, so they can't just let the public use them.
Re: Imagen, a text-to-image diffusion model
#398I apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself. Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are. Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication . And short of FAIR? D2 lacrosse. There are exceptions to such…
> Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication. This is ... very incorrect. I am very certain (95%+) that Google had nothing even close to GPT-3 at the time of its release. It's been 2 full years since GPT-3 was released, and even longer since OpenAI actually trained it. That's not to talk about any of the other things OpenAI/FAIR has released that were SOTA at the time…
Re: Imagen, a text-to-image diffusion model
#399Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.
“Oh our tech is so dangerous and amazing it could turn the world upside down” yet we hand it to random Bluechecks on Twitter.
It’s just marketing
Re: Imagen, a text-to-image diffusion model
#400Earlier quoted context omitted.
>Could skewing search results, i.e. hiding the bias of the real world Your logic seems to rest on this assumption which I don't think is justified. "Skewing search results" is not the same as "hiding the biases of the real world". Showing the most statistically likely result is not the same as showing the world how it truly is. A generic nurse is statistically going to be female most of the time. However, a model tha…
> I think it is reasonable for the creators to avoid sharing models known to not be smart enough to avoid exaggerating real world biases. Every model will have some random biases. Some of those random biases will undesirably exaggerate the real world. Every model will undesirably exaggerate something. Therefore no model should be shared. Your goal is nice, but impractical?
I said "It is reasonable... to avoid sharing models". That is an acknowledged that the creators are acting reasonably. It does not imply anything as extreme as "no model should be shared". The only way to get from A to B there is for you to assume that I think there is only one reasonable response and every other possible reaction is unreasonable. Doesn't that seem like a silly assumption?