Live data from Hacker News

DeepFloyd IF: open-source text-to-image model

github.com

71–80 of 237 posts

Re: DeepFloyd IF: open-source text-to-image model

#71
post #69

Earlier quoted context omitted.

Compositional power might mean "the image more resembles the composition you want and describe" i.e. if you say "a red cube on a green sphere" in DeepFloyd, you will get it. If you say that in MJ, you won't. That means you have more power to compose the image you want with this tool.

No, it does mean you don't understand how to prompt MJ, you don't understand it's language. You might like french more but it doesn't mean that it's a better language than english. MJ even says that their model doesn't understand language like humans do in their FAQs...

The point of text to image model is for them to accept natural language (yes, in practice, they all benefit from specialized prompting done with an understanding of model quirks, but that’s not the goal.)

Re: DeepFloyd IF: open-source text-to-image model

#72
post #69

Earlier quoted context omitted.

Compositional power might mean "the image more resembles the composition you want and describe" i.e. if you say "a red cube on a green sphere" in DeepFloyd, you will get it. If you say that in MJ, you won't. That means you have more power to compose the image you want with this tool.

No, it does mean you don't understand how to prompt MJ, you don't understand it's language. You might like french more but it doesn't mean that it's a better language than english. MJ even says that their model doesn't understand language like humans do in their FAQs...

Please give me a prompt which would back up your claim that MJ can do this! I'd love to learn

Re: DeepFloyd IF: open-source text-to-image model

#73
post #69

Earlier quoted context omitted.

Compositional power might mean "the image more resembles the composition you want and describe" i.e. if you say "a red cube on a green sphere" in DeepFloyd, you will get it. If you say that in MJ, you won't. That means you have more power to compose the image you want with this tool.

No, it does mean you don't understand how to prompt MJ, you don't understand it's language. You might like french more but it doesn't mean that it's a better language than english. MJ even says that their model doesn't understand language like humans do in their FAQs...

For example, please help me understand how to help MJ do the prompt above!

https://twitter.com/eb_french/status/1651365078137200640

BTW I am a HUGE fan of MJ, and attend the office hours, and have done 35k+ images there. So you may have misinterpreted how much of a supporter of it I am.

Re: DeepFloyd IF: open-source text-to-image model

#74

Earlier quoted context omitted.

Once these are quantized (I assume they can be), they should be ~1/4th the size. Can anyone explain why it needs so much ram in the first place though? 4.3B is only ~9GB at 16bit (I'm not as familiar with image models). I'm really happy to see that fits under 24GB - that's what I consider the limit for being able to run on "consumer hardware".

They took down the blogpost, but from what I remember the model is composite and consists of a text encoder as well as 3 "stages": 1. (11B) T5-XXL text encoder [1] 2. (4.3B) Stage 1 UNet 3. (1.3B) Stage 2 upscaler (64x64 -> 256x256) 4. (?B) Stage 3 upscaler (256x256 -> 1024x1024) Resolution numbers could be off though. Also the third stage can apparently use the existing stable diffusion x4, or a new upscaler that th…

The entire T5-XXL model is 11B but you don't need the decoder.

Re: DeepFloyd IF: open-source text-to-image model

#75
post #69

Earlier quoted context omitted.

No, it does mean you don't understand how to prompt MJ, you don't understand it's language. You might like french more but it doesn't mean that it's a better language than english. MJ even says that their model doesn't understand language like humans do in their FAQs...

For example, please help me understand how to help MJ do the prompt above! https://twitter.com/eb_french/status/1651365078137200640 BTW I am a HUGE fan of MJ, and attend the office hours, and have done 35k+ images there. So you may have misinterpreted how much of a supporter of it I am.

first shot but it's a starting point. With 35k images you should be able to do this yourself.

a single (green sphere), with a single (red cube), balancing ((on top))

https://imgur.com/a/QePgM6I

Re: DeepFloyd IF: open-source text-to-image model

#76
post #69

Earlier quoted context omitted.

No, it does mean you don't understand how to prompt MJ, you don't understand it's language. You might like french more but it doesn't mean that it's a better language than english. MJ even says that their model doesn't understand language like humans do in their FAQs...

The point of text to image model is for them to accept natural language (yes, in practice, they all benefit from specialized prompting done with an understanding of model quirks, but that’s not the goal.)

The way to prompt is a preference like programming languages ofc the layman might use the javascript of generative models because it's easier to start and there are a lot of tutorials but some might prefer something more exoctic which can produce the same or better quality. Whatever floats your boat but don't try to compare it like the guy in OPs tweets.

MJ and stable also make clear that their models don't understand language like humans do.

Re: DeepFloyd IF: open-source text-to-image model

#77
post #75

Earlier quoted context omitted.

For example, please help me understand how to help MJ do the prompt above! https://twitter.com/eb_french/status/1651365078137200640 BTW I am a HUGE fan of MJ, and attend the office hours, and have done 35k+ images there. So you may have misinterpreted how much of a supporter of it I am.

first shot but it's a starting point. With 35k images you should be able to do this yourself. a single (green sphere), with a single (red cube), balancing ((on top)) https://imgur.com/a/QePgM6I

Interesting, I do not get the results you do. What additional parameters are you using? Here is a link to some of my tests, with all default settings, some in v5 some in v4. https://twitter.com/eb_french/status/1651370091869786112

0/16 images have a red cube on a green sphere.

Re: DeepFloyd IF: open-source text-to-image model

#78
post #75

Earlier quoted context omitted.

For example, please help me understand how to help MJ do the prompt above! https://twitter.com/eb_french/status/1651365078137200640 BTW I am a HUGE fan of MJ, and attend the office hours, and have done 35k+ images there. So you may have misinterpreted how much of a supporter of it I am.

first shot but it's a starting point. With 35k images you should be able to do this yourself. a single (green sphere), with a single (red cube), balancing ((on top)) https://imgur.com/a/QePgM6I

Here are 16 more images with set seeds this time. Could you provide a complete prompt & seed to generate an image where MJ does this well? https://twitter.com/eb_french/status/1651371581514579969

Re: DeepFloyd IF: open-source text-to-image model

#79
post #75

Earlier quoted context omitted.

first shot but it's a starting point. With 35k images you should be able to do this yourself. a single (green sphere), with a single (red cube), balancing ((on top)) https://imgur.com/a/QePgM6I

Interesting, I do not get the results you do. What additional parameters are you using? Here is a link to some of my tests, with all default settings, some in v5 some in v4. https://twitter.com/eb_french/status/1651370091869786112 0/16 images have a red cube on a green sphere.

none and as an experienced user you should know that's it's not one shot and most of the time not even few shot... You can't compare cherry picked press images with few shots of a 5 second prompt. I don't know why you want to hype something up if you can't really compare it. It seems extremly attention grifting.

Just look at their cherry picks in this discord... https://discord.com/invite/pxewcvSvNx . It's overfitted on images with copyright (afghan girl) and doesn't show more "compositional power" at all most of the time ignoring half of the prompt.

Re: DeepFloyd IF: open-source text-to-image model

#80

Earlier quoted context omitted.

That's already prohibited by, you know, those very same copyright and privacy laws. Adding those same prohibitions to the license not only makes the software nonfree, but pointlessly does so.

Its not pointless, it means the model licensor has a claim against you, as well as whoever would for violating the referenced laws; it also means, and this is probably more important, that in some juridictions, the model licensor has a better defense against liability for contributory infringement if the licensee infringes. EDIT: That said, it’s unambiguously not open source.

> it means the model licensor has a claim against you

Right, but to what end? The only reason the licensor should care one way or another is the licensor being held liable for what folks do with the software, in which case...

> it also means, and this is probably more important, that in some juridictions, the model licensor has a better defense against liability for contributory infringement if the licensee infringes.

Do hardware stores need to demand "thou shalt not use this tool to kill people" to their customers to avoid liability for axe murders under such jurisdictions? Or car manufacturers needing to specify "you will not use this product to run over schoolchildren at crosswalks"?

Like, I'm sure such jurisdictions exist, but I somehow doubt license terms in an EULA nobody (except for us nerds) will ever read would be sufficient in such a kangaroo court.

(EDIT: also, I'm pretty sure the standard warranty disclaimer in your average FOSS license already covers this, without making the software nonfree in the process)

Post reply on HN