Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

571–580 of 661 posts

Re: Imagen, a text-to-image diffusion model

#571
post #42

Earlier quoted context omitted.

This raises some really interesting questions. We certainly don't want to perpetuate harmful stereotypes. But is it a flaw that the model encodes the world as it really is, statistically, rather than as we would like it to be? By this I mean that there are more light-skinned people in the west than dark, and there are more women nurses than men, which is reflected in the model's training data. If the model only gener…

It depends on whether you'd like the model to learn casual or correlative relationships. If you want the model to understand what a "nurse" actually is, then it shouldn't be associated with female. If you want the model to understand how the word "nurse" is usually used, without regard for what a "nurse" actually is, then associating it with female is fine. The issue with a correlative model is that it can easily be…

The meaning of the word "nurse" is determined by how the word "nurse" is used and understood.

Perhaps what "nurse" means isn't what "nurse" should mean, but what people mean when they say "nurse" is what "nurse" means.

Re: Imagen, a text-to-image diffusion model

#572
post #282

Earlier quoted context omitted.

> Humans overwhelmingly learn meaning by use, not by definition Preliminarily and provisionally. Then, they start discussing their concepts - it is the very definition of Intelligence.

Most humans don’t do that for most things they have a notion of in their head. It would be much too time consuming to start discussing the meaning of even just a significant fraction of them. For a rough reference point, the English language has over 150.000 words that you could each discuss the meaning of and try to come up with a definition. Not to speak of the difficulties to make that set of definitions noncircul…

(Mental entities are very many more than the hundred thousand, out of composition, cartesianity etc. So-called "protocols" (after logical positivism) are part of them, relating more entities with space and time. Also, by speaking of "circular definitions" you are, like others, confusing mental definitions with formal definitions.)

So? Draw your consequences.

Following what was said, you are stating that "a staggering large number of people are unintelligent". Well, ok, that was noted. Scolio: if unintelligent, they should refrain from expressing judgement (you are really stating their non-judgement), why all the actual expression? If unintelligent actors, they are liabilities, why this overwhelming employment in the job market?

Thing is, as unintelligent as you depict them quantitatively, the internal processing that constitutes intelligence proceeds in many even when scarce, even when choked by some counterproductive bad formation - processing is the natural functioning. And then, the right Paretian side will "do the job" that the vast remainder will not do, and process notions actively (more, "encouragingly" - the process is importantly unconscious, many low-level layers are) and proficiently.

And the very Paretian prospect will reveal, there will be a number of shallow takes, largely shared, on some idea, and other intensively more refined takes, more rare, on the same idea. That shows you a distinction between "use" and the asymptotic approximation to meanings as achieved by intellectual application.

Re: Imagen, a text-to-image diffusion model

#573
post #493

Earlier quoted context omitted.

People training newer models just have to look for the "Imagen" tag or the Dall-E2 rainbow at the corner and heuristically exclude images having these. This is trivial. Unless you assume there are bad actors who will crop out the tags. Not many people now have access to Dall-E2 or will have access to Imagen. As someone working in Vision, I am also thinking about whether to include such images deliberately. Using imag…

Most images you see from these services will not have a watermark on them. Cropping is trivial.

It's ironic, seeing people who build models trained on other people's work (which is in no way credited) to be worried about origin and credit.

Re: Imagen, a text-to-image diffusion model

#574
post #12

Interesting discovery they made > We show that scaling the pretrained text encoder size is more important than scaling the diffusion model size. There seems to be an unexpected level of synergy between text and vision models. Can't wait to see what video and audio modalities will add to the mix.

Basically makes sense, no? DALLE-2 suffered from misunderstanding propositional logic, treating prompts as less structured then it should have. That's a text model issue! Compared to that, scaling up the image isn't as important (especially with a few passes).

Is there a way to confirm that this extra processing relates to the language structure, and not the processing of concepts?

I wouldn’t be surprised if the lack of video and 3D understanding in the image dataset training fails to understand things like the fear of heights, and the concept of gravity ends up being learned in the text processing weights.

Re: Imagen, a text-to-image diffusion model

#575

Earlier quoted context omitted.

It’s just my opinion but I think the meme you’re talking about is deeply related to other branches of science and philosophy: ranging from the trust old saw about AI being anything a computer hasn’t done yet to deep meditations on the nature of consciousness. They’re all fundamentally anthropocentric: people argue until they are blue in the face about what “intelligent” means but it’s always implicit that what they r…

I'd argue that there is probably at least one leap in terms of human-level writing which isn't just pure prediction. Humans write with intent , which is how we can maintain long run structure. I definitely write like GPT while I'm not paying attention, but with the executive on the task I outperform it. For all we know this is solvable with some small tweak to architecture, and I rather doubt that a model which has s…

I agree that intent is the missing piece so far. GTP can respond better to prompts than most people, but does so with a complete lack of intent. The human provides 100% of it.

Re: Imagen, a text-to-image diffusion model

#576
post #462

Earlier quoted context omitted.

I agree. How cool would it be to get an 8 min version of your favorite song? Or an instant DnB remix? Or 10 more songs in the style of your favorite album?

You can sort of do that with https://fairuseify.ml

I believe that this tech is possible, but this site doesn't provide it. Look at the source of the page: it's just a bunch of sleeps and then you 'download' the same file you provided.

Re: Imagen, a text-to-image diffusion model

#577
post #230

Earlier quoted context omitted.

Revert back to average representation of a nurse (give no weight to unspecified criterias, gender, age, skin-color, religion, country, hair-style, no style whether it's a drawing or a photography, no information about the year it was made, etc). “hey artist, draw me a nurse.” “Hmm okay, do you want it a guy or girl?” “Don’t ask me, just draw what I’m saying.” - Ok, I'll draw you what an average nurse looks like. - Wa…

The average nurse has three-halfs of a tit.

Is it not incredible that after so many decades talking about local minima there is now some supposition that all of them must merge?

Re: Imagen, a text-to-image diffusion model

#578
post #432

Off topic, but this caught my attention: “In future work we will explore a framework for responsible externalization that balances the value of external auditing with the risks of unrestricted open-access.” I work for a big org myself, and I’ve wondered what it is exactly that makes people in big orgs so bad at saying things.

I think they were being careful not to be too quotable there on CNN.

Re: Imagen, a text-to-image diffusion model

#579

Interesting to me that this one can draw legible text. DALLE models seem to generate weird glyphs that only look like text. The examples they show here have perfectly legible characters and correct spelling. The difference between this and DALLE makes me suspicious / curious. I wish I could play with this model.

DALLE1 was able to render text[0]. That DALLE2 isn't probably is a tradeoff introduced by unCLIP in exchange for diverse results. Now the google model is better yet and doesn't have to make that tradeoff.

[0] https://openai.com/blog/dall-e/#text-rendering

Re: Imagen, a text-to-image diffusion model

#580
post #129

All of these AI findings are cool in theory. But until its accessible to some decent amount of people/customers - its basically useless fluff. You can tell me those pictures are generated by an AI and I might believe it, but until real people can actually test it... it's easy enough to fake. This page isn't even the remotest bit legit by the URL, It looks nicely put together and that's about it. Could have easily put…

Inference times are key. If it can't be produced within reasonable latency, then there will be no real world use case for it because it's simply too expensive to run inference at scale.

There's been much prior work done to take these models down from datacenter size to single GPU size. Given continued work in that area and improving GPU performance it seems like it's just a matter of years before inference can be cheap and local for even the most impressive of generation.
Post reply on HN