Live data from Hacker News

DeepFloyd IF: open-source text-to-image model

github.com

101–110 of 237 posts

Re: DeepFloyd IF: open-source text-to-image model

#101

Earlier quoted context omitted.

Midjourney always look very aesthetic pleasing, I guess because of their RLHF tuning with Discord data... But it doesn't really follow prompts as well as Dall-e for example. But in the end, people want pretty pictures. So is a complicated situation.

The tweet they shared is from February and uses an outdated version of MJ, this is what I got from V5: https://i.imgur.com/0uxtZDe.png Midjourney does much better overall . Composition is neat, but MJ is so incredibly far ahead in terms of quality of output, it honestly doesn't matter if you have to go and do composition manually (and with new AI based tools, that's easier than ever too. Do a bad cut and paste job th…

But it didn't work? In yours there is no "Nexus", no smiling, no frowning, and man on right doesn't look Asian? Compared with image in Tweet, MJ failed at this task.

Re: DeepFloyd IF: open-source text-to-image model

#102

Earlier quoted context omitted.

There's fundamental tradeoffs, as there always will be when you're compressing things into an image model. So, this is going to have new different issues. Since it's similar to Imagen, it probably can't handle long complex prompts as well, since they developed Parti afterward. Here's my question: are there any image models where, if you prompt "1+1", you get an image showing "3"?

> So, this is going to have new different issues. Well, yeah, its a bigger set of models (particular the language model) that takes more resources (both to train and for inference.) That’s the tradeoff. > Here’s my question: are there any image models where, if you prompt “1+1”, you get an image showing “3”? You want a t2i model that does arithmetic in the prompt, translates to it to “text displaying the number ”, bu…

Pretend I wrote 2, edit timeout closed.

"Actually thinking about your prompt" is a necessary part of being able to make the prompts natural language instead of a long list of fantasy google image search terms.

Useful example being "my bedroom but in a new color", but some things I've typed into Midjourney that don't work include "a really long guinea pig" (you get a regular size one), "world's best coffee" (the coffee cup gets a world on it), etc. It's just too literal.

And yes, preprocessing with an LLM could do this.

Re: DeepFloyd IF: open-source text-to-image model

#103
post #79

Earlier quoted context omitted.

none and as an experienced user you should know that's it's not one shot and most of the time not even few shot... You can't compare cherry picked press images with few shots of a 5 second prompt. I don't know why you want to hype something up if you can't really compare it. It seems extremly attention grifting. Just look at their cherry picks in this discord... https://discord.com/invite/pxewcvSvNx . It's overfitted…

> as an experienced user you should know that's it's not one shot Being "not one shot" for most nontrivial prompts is a failure of current t2i models, its what they all strive for and its what DF supposedly does a lot better. And, while its possible to spin things pretty hard when people can't bang on it themselves, I think the indication is that it is, in fact, a major leap forward from the best current consumer-ava…

Are you associated with deepfloyd?

It's not a major leap not even a small one because it's exactly like imagen. It's stability giving some Ukrainian refugees compute time to train "their" model for publicity. It's about the whom and not what as it should be.

I "feel" nothing I am telling it how it is. Look at the afghan girl example again it. Close up portrait, same clothing, same comp, expressive eyes... and most important burn in like every other overfitted image in diffusion networks.

You guys all want it to be something special and I get it, new content, new shiny toy but it's neither a good architecture nor a good implementation.

Re: DeepFloyd IF: open-source text-to-image model

#104

Example of how much better it can do compared to midjourney, on a complex prompt: https://twitter.com/eb_french/status/1623823175170805760 It is able to put people on the left/right and put the correct t-shirts and facial expressions on each one. This is compared to mj which just mixes together a soup of every word you use and plops it out into the image. Huge MJ fan of course, it's amazing, but having compositional…

Can't wait to see how all of this is going to look like in ten years. I know we are all nitpicking right now but these results are totally mindblowing already.

I wonder when we can start fuzzing brains. Wire you up to a machine that measures happiness or anxiety or anger or whatever, and keep re-generating results that hone in on the given emotion.

Re: DeepFloyd IF: open-source text-to-image model

#105

Earlier quoted context omitted.

There's fundamental tradeoffs, as there always will be when you're compressing things into an image model. So, this is going to have new different issues. Since it's similar to Imagen, it probably can't handle long complex prompts as well, since they developed Parti afterward. Here's my question: are there any image models where, if you prompt "1+1", you get an image showing "3"?

> So, this is going to have new different issues. Well, yeah, its a bigger set of models (particular the language model) that takes more resources (both to train and for inference.) That’s the tradeoff. > Here’s my question: are there any image models where, if you prompt “1+1”, you get an image showing “3”? You want a t2i model that does arithmetic in the prompt, translates to it to “text displaying the number ”, bu…

I don't think they're saying that's a goal, I think they're curious if it is the case. LLMs are bad at arithmetic, this uses a LLM to process the prompt, that class of result seems plausible.

Re: DeepFloyd IF: open-source text-to-image model

#106

Earlier quoted context omitted.

Midjourney always look very aesthetic pleasing, I guess because of their RLHF tuning with Discord data... But it doesn't really follow prompts as well as Dall-e for example. But in the end, people want pretty pictures. So is a complicated situation.

The tweet they shared is from February and uses an outdated version of MJ, this is what I got from V5: https://i.imgur.com/0uxtZDe.png Midjourney does much better overall . Composition is neat, but MJ is so incredibly far ahead in terms of quality of output, it honestly doesn't matter if you have to go and do composition manually (and with new AI based tools, that's easier than ever too. Do a bad cut and paste job th…

Well... Kind of, photobashing with midjourney doesn't guarantee you the same image or even necessarily the objects in the same places, even if you increase the image weight value up to its maximum of two. ('--iw 2')

Many times you'll have no other choice but to use a diffusion model with img2img.

I agree with OP though, the market has spoken and the vast majority of people use prompts hardly more nuanced than a 90s Mad Magazine book of Mad Libs.

Re: DeepFloyd IF: open-source text-to-image model

#107
post #90

Earlier quoted context omitted.

Stable Diffusion 2.x has been supported for a while.

Yup, lots of misinformation in this thread from those who are not in-the-know. Automatic1111 is the defacto main UI for these kind of models. It will be supported there, quite quickly.

Automatic hasn't been updated for several weeks at this point. Several people are trying to fork the repo to make their own continuation.

Re: DeepFloyd IF: open-source text-to-image model

#108
post #94
post #41

Earlier quoted context omitted.

It prohibits both commercial use, whether or not you break regional laws; and it prohibits breaking certain laws. As another user said, encoding the law into a licence is pointless but makes it non-free. There are also problematic restrictions on your ability to modify the software under clause 2(c). And nor do you have the right to sublicence, it's not clear to me what rights somebody has if you give them a copy.

Where does it prohibit commercial use? I’m not seeing that in the license.

The model license, 1(a) and 2(a)(i).

Re: DeepFloyd IF: open-source text-to-image model

#109

Earlier quoted context omitted.

Can't wait to see how all of this is going to look like in ten years. I know we are all nitpicking right now but these results are totally mindblowing already.

I wonder when we can start fuzzing brains. Wire you up to a machine that measures happiness or anxiety or anger or whatever, and keep re-generating results that hone in on the given emotion.

Sometimes you have to ask yourself if that's a world you really want to live in.

Re: DeepFloyd IF: open-source text-to-image model

#110

Example of how much better it can do compared to midjourney, on a complex prompt: https://twitter.com/eb_french/status/1623823175170805760 It is able to put people on the left/right and put the correct t-shirts and facial expressions on each one. This is compared to mj which just mixes together a soup of every word you use and plops it out into the image. Huge MJ fan of course, it's amazing, but having compositional…

Can't wait to see how all of this is going to look like in ten years. I know we are all nitpicking right now but these results are totally mindblowing already.

The thing is all these things can go downhill as well. All these things are cool today just like Google was a decade ago.
Post reply on HN