Earlier quoted context omitted.
Midjourney always look very aesthetic pleasing, I guess because of their RLHF tuning with Discord data... But it doesn't really follow prompts as well as Dall-e for example. But in the end, people want pretty pictures. So is a complicated situation.
The tweet they shared is from February and uses an outdated version of MJ, this is what I got from V5: https://i.imgur.com/0uxtZDe.png Midjourney does much better overall . Composition is neat, but MJ is so incredibly far ahead in terms of quality of output, it honestly doesn't matter if you have to go and do composition manually (and with new AI based tools, that's easier than ever too. Do a bad cut and paste job th…
DeepFloyd IF: open-source text-to-image model
101–110 of 237 posts
Re: DeepFloyd IF: open-source text-to-image model
#102Earlier quoted context omitted.
There's fundamental tradeoffs, as there always will be when you're compressing things into an image model. So, this is going to have new different issues. Since it's similar to Imagen, it probably can't handle long complex prompts as well, since they developed Parti afterward. Here's my question: are there any image models where, if you prompt "1+1", you get an image showing "3"?
> So, this is going to have new different issues. Well, yeah, its a bigger set of models (particular the language model) that takes more resources (both to train and for inference.) That’s the tradeoff. > Here’s my question: are there any image models where, if you prompt “1+1”, you get an image showing “3”? You want a t2i model that does arithmetic in the prompt, translates to it to “text displaying the number ”, bu…
"Actually thinking about your prompt" is a necessary part of being able to make the prompts natural language instead of a long list of fantasy google image search terms.
Useful example being "my bedroom but in a new color", but some things I've typed into Midjourney that don't work include "a really long guinea pig" (you get a regular size one), "world's best coffee" (the coffee cup gets a world on it), etc. It's just too literal.
And yes, preprocessing with an LLM could do this.
Re: DeepFloyd IF: open-source text-to-image model
#103Earlier quoted context omitted.
none and as an experienced user you should know that's it's not one shot and most of the time not even few shot... You can't compare cherry picked press images with few shots of a 5 second prompt. I don't know why you want to hype something up if you can't really compare it. It seems extremly attention grifting. Just look at their cherry picks in this discord... https://discord.com/invite/pxewcvSvNx . It's overfitted…
> as an experienced user you should know that's it's not one shot Being "not one shot" for most nontrivial prompts is a failure of current t2i models, its what they all strive for and its what DF supposedly does a lot better. And, while its possible to spin things pretty hard when people can't bang on it themselves, I think the indication is that it is, in fact, a major leap forward from the best current consumer-ava…
It's not a major leap not even a small one because it's exactly like imagen. It's stability giving some Ukrainian refugees compute time to train "their" model for publicity. It's about the whom and not what as it should be.
I "feel" nothing I am telling it how it is. Look at the afghan girl example again it. Close up portrait, same clothing, same comp, expressive eyes... and most important burn in like every other overfitted image in diffusion networks.
You guys all want it to be something special and I get it, new content, new shiny toy but it's neither a good architecture nor a good implementation.
Re: DeepFloyd IF: open-source text-to-image model
#104Example of how much better it can do compared to midjourney, on a complex prompt: https://twitter.com/eb_french/status/1623823175170805760 It is able to put people on the left/right and put the correct t-shirts and facial expressions on each one. This is compared to mj which just mixes together a soup of every word you use and plops it out into the image. Huge MJ fan of course, it's amazing, but having compositional…
Can't wait to see how all of this is going to look like in ten years. I know we are all nitpicking right now but these results are totally mindblowing already.
Re: DeepFloyd IF: open-source text-to-image model
#105Earlier quoted context omitted.
There's fundamental tradeoffs, as there always will be when you're compressing things into an image model. So, this is going to have new different issues. Since it's similar to Imagen, it probably can't handle long complex prompts as well, since they developed Parti afterward. Here's my question: are there any image models where, if you prompt "1+1", you get an image showing "3"?
> So, this is going to have new different issues. Well, yeah, its a bigger set of models (particular the language model) that takes more resources (both to train and for inference.) That’s the tradeoff. > Here’s my question: are there any image models where, if you prompt “1+1”, you get an image showing “3”? You want a t2i model that does arithmetic in the prompt, translates to it to “text displaying the number ”, bu…
Re: DeepFloyd IF: open-source text-to-image model
#106Earlier quoted context omitted.
Midjourney always look very aesthetic pleasing, I guess because of their RLHF tuning with Discord data... But it doesn't really follow prompts as well as Dall-e for example. But in the end, people want pretty pictures. So is a complicated situation.
The tweet they shared is from February and uses an outdated version of MJ, this is what I got from V5: https://i.imgur.com/0uxtZDe.png Midjourney does much better overall . Composition is neat, but MJ is so incredibly far ahead in terms of quality of output, it honestly doesn't matter if you have to go and do composition manually (and with new AI based tools, that's easier than ever too. Do a bad cut and paste job th…
Many times you'll have no other choice but to use a diffusion model with img2img.
I agree with OP though, the market has spoken and the vast majority of people use prompts hardly more nuanced than a 90s Mad Magazine book of Mad Libs.
Re: DeepFloyd IF: open-source text-to-image model
#107Earlier quoted context omitted.
Stable Diffusion 2.x has been supported for a while.
Yup, lots of misinformation in this thread from those who are not in-the-know. Automatic1111 is the defacto main UI for these kind of models. It will be supported there, quite quickly.
Re: DeepFloyd IF: open-source text-to-image model
#108Earlier quoted context omitted.
It prohibits both commercial use, whether or not you break regional laws; and it prohibits breaking certain laws. As another user said, encoding the law into a licence is pointless but makes it non-free. There are also problematic restrictions on your ability to modify the software under clause 2(c). And nor do you have the right to sublicence, it's not clear to me what rights somebody has if you give them a copy.
Where does it prohibit commercial use? I’m not seeing that in the license.
Re: DeepFloyd IF: open-source text-to-image model
#109Earlier quoted context omitted.
Can't wait to see how all of this is going to look like in ten years. I know we are all nitpicking right now but these results are totally mindblowing already.
I wonder when we can start fuzzing brains. Wire you up to a machine that measures happiness or anxiety or anger or whatever, and keep re-generating results that hone in on the given emotion.
Re: DeepFloyd IF: open-source text-to-image model
#110Example of how much better it can do compared to midjourney, on a complex prompt: https://twitter.com/eb_french/status/1623823175170805760 It is able to put people on the left/right and put the correct t-shirts and facial expressions on each one. This is compared to mj which just mixes together a soup of every word you use and plops it out into the image. Huge MJ fan of course, it's amazing, but having compositional…
Can't wait to see how all of this is going to look like in ten years. I know we are all nitpicking right now but these results are totally mindblowing already.