Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

431–440 of 661 posts

Re: Imagen, a text-to-image diffusion model

#431
post #398

Earlier quoted context omitted.

> Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication. This is ... very incorrect. I am very certain (95%+) that Google had nothing even close to GPT-3 at the time of its release. It's been 2 full years since GPT-3 was released, and even longer since OpenAI actually trained it. That's not to talk about any of the other things OpenAI/FAIR has released that were SOTA at the time…

Yeah, at the time, GB was still very big on mixture-of-expert models and bidirectional models like T5. (I'm not too enthusiastic about the former, but the latter has been a great model family and even if not GPT-3, still awesome.) DeepMind pivoted faster than GB, based on Gopher's reported training date, and GB followed some time after. But definitely neither had their own GPT-3-scale dense Transformer when GPT-3 was…

At the risk of sounding like I’m trying to defend a position that I’ve already conceded is an oversimplification, I’m frankly a little skeptical of how we can even know that.

GPT is, opaque. It’s somewhere between common knowledge and conspiracy theory that it gets a helping hand from Turks when it gets in over its head.

The exact details of why a BERT-style transformer, or any of the zillion other lookalikes, isn’t just over-fitting Wikipedia the more corpus and compute you feed to its insatiable maw has always seemed a little big on claims and light on reproducibility.

I don’t think there are many attention skeptics in language modeling, it’s a good idea that you can demo on a gaming PC. Transformers demonstrably work, and a better beam-search (or whatever) hits the armchair Turing test harder for a given compute budget.

But having seen some of this stuff play out at scale, and admittedly this is purely anecdotal, these things are basically asking the question: “if I overfit all human language on the Internet, is that a bad thing?”

It’s my personal suspicion that this is the dominant term, and it’s my personal belief that Google’s ability to do both corpus and model parallelism at Jeff Dean levels while simultaneously building out hardware to the exact precision required is unique by a long way.

But, to be more accurate than I was in my original comment, I don’t know most of that in the sense that would be required by peer-review, let alone a jury. It’s just an educated guess.

Re: Imagen, a text-to-image diffusion model

#432
Off topic, but this caught my attention:

“In future work we will explore a framework for responsible externalization that balances the value of external auditing with the risks of unrestricted open-access.”

I work for a big org myself, and I’ve wondered what it is exactly that makes people in big orgs so bad at saying things.

Re: Imagen, a text-to-image diffusion model

#433

Earlier quoted context omitted.

The AI ethics thing is just a PR larp at this point. “Oh our tech is so dangerous and amazing it could turn the world upside down” yet we hand it to random Bluechecks on Twitter. It’s just marketing

You know Twitter and Google are different companies, right?

The commenter was probably referring to the fact that the people who tend to get access to things like GPT3 or DALL-E 2 tend to be people with large Twitter followings (and blue checks), and that there may be a significant marketing component to this fact.

Re: Imagen, a text-to-image diffusion model

#435
post #370

Interesting to me that this one can draw legible text. DALLE models seem to generate weird glyphs that only look like text. The examples they show here have perfectly legible characters and correct spelling. The difference between this and DALLE makes me suspicious / curious. I wish I could play with this model.

Imagen takes text embeddings, OpenAI model takes image embeddings instead, this is the reason. There are other models that can generate text: latent diffusion trained on LAION-400M, GLIDE, DALL-E (1).

My understanding of the terms text and image embeddings is that they are ways of representing text or images as vectors. But, I don't understand how that would help with the process of actually drawing the symbols for those letters.

Re: Imagen, a text-to-image diffusion model

#437
I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely available the internet will become harder to scrape for good data and models will start training on their own output. The internet will also probably get worse for humans since search results will be completely polluted with these "sort of realistic" images which can ultimately be spit out at breakneck speed by smashing words from a dictionary together...

Re: Imagen, a text-to-image diffusion model

#439

I thought I was doing well after not being overly surprised by DALL-E 2 or Gato. How am I still not calibrated on this stuff? I know I am meant to be the one who constantly argues that language models already have sophisticated semantic understanding, and that you don't need visual senses to learn grounded world knowledge of this sort, but come on, you don't get to just throw T5 in a multimodal model as-is and have i…

It’s just my opinion but I think the meme you’re talking about is deeply related to other branches of science and philosophy: ranging from the trust old saw about AI being anything a computer hasn’t done yet to deep meditations on the nature of consciousness. They’re all fundamentally anthropocentric: people argue until they are blue in the face about what “intelligent” means but it’s always implicit that what they r…

I'd argue that there is probably at least one leap in terms of human-level writing which isn't just pure prediction. Humans write with intent, which is how we can maintain long run structure. I definitely write like GPT while I'm not paying attention, but with the executive on the task I outperform it. For all we know this is solvable with some small tweak to architecture, and I rather doubt that a model which has solved this problem need be conscious (though our own solution seems correlated with consciousness), but it is one more step.

Re: Imagen, a text-to-image diffusion model

#440
post #132

Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.

The AI ethics thing is just a PR larp at this point. “Oh our tech is so dangerous and amazing it could turn the world upside down” yet we hand it to random Bluechecks on Twitter. It’s just marketing

And that's why we need to be shutting this stuff down, completely.

Someone tried to say there were ethics committees etc the other day...what a bad joke. Who checks the ethics committee is making ethical decisions?

I was told I "didn't know what" I was talking about, excuse from some over-important know-it-all who didn't know what ethics was, i.e. they don't know what they are talking about.

Post reply on HN