How the fck are things advancing so fast? Is it about to level off …or extend to new domains? What’s a comparable set of technical advances?
Imagen, a text-to-image diffusion model
321–330 of 661 posts
Re: Imagen, a text-to-image diffusion model
#322Earlier quoted context omitted.
The big labs have become very sensitive with large model releases. It's too easy to make them generate bad PR, to the point of not releasing almost any of them. Flamingo was also a pretty great vison-language model that wasn't released, not even in a demo. PaLM is supposedly better than GPT-3 but closed off. It will probably take a year for open source models to appear.
The largest models which generate the headline benchmarks are never released after any number of years, it seems. Very difficult to replicate results.
Re: Imagen, a text-to-image diffusion model
#323Earlier quoted context omitted.
It's an unfortunate reflection of reality. There are three possible outcomes: 1. The model provides a reflection of reality, as politically inconvenient and hurtful as it may be. 2. The model provides an intentionally obfuscated version with either random traits or non correlative traits. 3. The model refuses to answer. Which of these is ideal to you?
What makes you think those are the only options? Why can't we have an option that the model returns a range of different outputs based off a prompt? A model that returns 100% of nurses as female might be statistically more accurate than a model that returns 50% of nurses as female, but it is still not an accurate reflection of the real world. I agree that the model shouldn't return a male nurse 50% of the time. Yet a…
Re: Imagen, a text-to-image diffusion model
#324How the fck are things advancing so fast? Is it about to level off …or extend to new domains? What’s a comparable set of technical advances?
An impressive advance would be a small model that’s capable of working from an external memory rather than memorizing it.
Re: Imagen, a text-to-image diffusion model
#325Earlier quoted context omitted.
Figure A.4 in the linked paper is a good high level overview of this model. Shame it was hidden away on page 19 in the appendix! Each box you see there has a section in the paper explaining it in more detail.
Uhh, yeah, I'm going to need much more of an ELI5 than that! Looking at Figure A.4, I understand (again, at a very high-level) the first step of "Frozen Text Encoder", and I have a decent understanding of the upsampling techniques used in the last 2 diffusion model steps, but the middle "Text-to-Image Diffusion Model" step that magically outputs a 64x64 pixel image of an actual golden retriever wearing an actual blue…
Re: Imagen, a text-to-image diffusion model
#326Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.
Re: Imagen, a text-to-image diffusion model
#327I apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself. Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are. Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication . And short of FAIR? D2 lacrosse. There are exceptions to such…
In short, it’s all about money.
For example: the high-frequency trading industry is estimated to have made somewhere between 2-3 billion dollars in all of 2020, profit/earnings. That’s a good weekend at Google.
HFT shops pay well, but not much different to top performers at FAANG.
People work in HFT because without taking a pay cut they can play real ball: they want to try themselves against the best.
Heavy learning people are no different in wanting both a competitive TC but maybe even more to be where the action is.
That’s currently Blade Runner Industries Ltd, but that could change.
Re: Imagen, a text-to-image diffusion model
#328Earlier quoted context omitted.
Google is very conservative about anything that can generate open-ended outputs. Also these models are still very expensive computationally.
They're expensive to train, but not awfully expensive to use. Especially if you have hundreds of images you want to generate (due to the way compute devices tend to get much more efficiency with a large batch size). Google could totally afford it, especially if the feature was hidden behind a button the user had to click, and not just run for every image search.
Re: Imagen, a text-to-image diffusion model
#329I apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself. Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are. Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication . And short of FAIR? D2 lacrosse. There are exceptions to such…
Google clearly demonstrates their unrivaled capability to leverage massive quantities of data and compute, but it’s premature to declare that they’ve secured victory in the AI Wars.
Re: Imagen, a text-to-image diffusion model
#330Earlier quoted context omitted.
Yes, the idea is that just because it doesn't align to Western ideals of what seems unbiased doesn't mean that the same is necessarily true for other cultures, and by failing to release the model because it doesn't conform to Western, left wing cultural expectations, the authors are ignoring the diversity of cultures that exist globally.
No, it's coming from a perspective of moral realism. It's an objective moral truth that racial and ethnic biases are bad. Yet most cultures around the world are racist to at least some degree, and to they extent that the cultures do, they are bad. The argument you're making, paraphrased, is that the idea that biases are bad is itself situated in particular cultural norms. While that is true to some degree, from a mor…
> from a moral realist perspective we can still objectively judge those cultural norms to be better or worse than alternatives
No, because depending on what set of values you have, it is easy to say that one set of biases is better than another. The entire point is that it should not be Google's role to make that judgement - people should be able to do it for themselves.