Live data from Hacker News

DeepFloyd IF: open-source text-to-image model

github.com

111–120 of 237 posts

Re: DeepFloyd IF: open-source text-to-image model

#111

Earlier quoted context omitted.

The tweet they shared is from February and uses an outdated version of MJ, this is what I got from V5: https://i.imgur.com/0uxtZDe.png Midjourney does much better overall . Composition is neat, but MJ is so incredibly far ahead in terms of quality of output, it honestly doesn't matter if you have to go and do composition manually (and with new AI based tools, that's easier than ever too. Do a bad cut and paste job th…

But it didn't work? In yours there is no "Nexus", no smiling, no frowning, and man on right doesn't look Asian? Compared with image in Tweet, MJ failed at this task.

That depends on what your goal was? If your goal was to get an AI model to generate copyrighted images, and misunderstand the relationship between Indians and Asians, then sure MJ failed (and I'm guessing that's the goal of the prompt).

But if I actually wanted a useful picture, I could work with what MJ gave me despite having minimal image editing skills. The DeepFloyd result looks like it's a 8-12 months behind what MJ gave and wouldn't be salvageable.

Re: DeepFloyd IF: open-source text-to-image model

#112

Earlier quoted context omitted.

The tweet they shared is from February and uses an outdated version of MJ, this is what I got from V5: https://i.imgur.com/0uxtZDe.png Midjourney does much better overall . Composition is neat, but MJ is so incredibly far ahead in terms of quality of output, it honestly doesn't matter if you have to go and do composition manually (and with new AI based tools, that's easier than ever too. Do a bad cut and paste job th…

But it didn't work? In yours there is no "Nexus", no smiling, no frowning, and man on right doesn't look Asian? Compared with image in Tweet, MJ failed at this task.

In what way does the man on the right not look like he could be from the absolutely enormous continent known as Asia?

Re: DeepFloyd IF: open-source text-to-image model

#113

Earlier quoted context omitted.

The tweet they shared is from February and uses an outdated version of MJ, this is what I got from V5: https://i.imgur.com/0uxtZDe.png Midjourney does much better overall . Composition is neat, but MJ is so incredibly far ahead in terms of quality of output, it honestly doesn't matter if you have to go and do composition manually (and with new AI based tools, that's easier than ever too. Do a bad cut and paste job th…

Well... Kind of, photobashing with midjourney doesn't guarantee you the same image or even necessarily the objects in the same places, even if you increase the image weight value up to its maximum of two. ('--iw 2') Many times you'll have no other choice but to use a diffusion model with img2img. I agree with OP though, the market has spoken and the vast majority of people use prompts hardly more nuanced than a 90s M…

I wasn't referring to photobashing, I meant firing up SD and ArtStudio for 10 minutes and getting something that looks amazing and has the desired composition.

Overall this feels like trying to get ChatGPT to do math: just let ChatGPT offload math to Wolfram.

Similarly I'd rather just offload the composition. Now we even have SAM which will happily pick out the parts of the image you want to compose

Re: DeepFloyd IF: open-source text-to-image model

#114

Earlier quoted context omitted.

Can't wait to see how all of this is going to look like in ten years. I know we are all nitpicking right now but these results are totally mindblowing already.

I wonder when we can start fuzzing brains. Wire you up to a machine that measures happiness or anxiety or anger or whatever, and keep re-generating results that hone in on the given emotion.

It's not because something can be done that it must be done.

Re: DeepFloyd IF: open-source text-to-image model

#115
post #55
post #27

Earlier quoted context omitted.

They are technically open source. It's just that the model license prohibits commercial use and the code license prohibits bypassing the filters. So it's kind of worse than closed source in a way because it's like a tease. With no API apparently. Theoretically large companies or rich people might be able to make a licensing agreement.

I am a lawyer, and as flimsy and wishy-washy as the term "open-source" already is, I can't even fathom what is meant by "open source" here? Are people suggesting that "look at the code but don't touch" actually fits what some people think of as open source?

NOT a lawyer here, basially they meant, here the thing, please don't sue me if you messed up.

Re: DeepFloyd IF: open-source text-to-image model

#116

Earlier quoted context omitted.

Can't wait to see how all of this is going to look like in ten years. I know we are all nitpicking right now but these results are totally mindblowing already.

I wonder when we can start fuzzing brains. Wire you up to a machine that measures happiness or anxiety or anger or whatever, and keep re-generating results that hone in on the given emotion.

So… heroin, cocaine, and methamphetamine, in memetic form?

I really hope that isn't actually possible.

Re: DeepFloyd IF: open-source text-to-image model

#117
post #76

Earlier quoted context omitted.

The point of text to image model is for them to accept natural language (yes, in practice, they all benefit from specialized prompting done with an understanding of model quirks, but that’s not the goal.)

The way to prompt is a preference like programming languages ofc the layman might use the javascript of generative models because it's easier to start and there are a lot of tutorials but some might prefer something more exoctic which can produce the same or better quality. Whatever floats your boat but don't try to compare it like the guy in OPs tweets. MJ and stable also make clear that their models don't understan…

> MJ and stable also make clear that their models don't understand language like humans do.

I believe the claim here is "and we would like them to".

Re: DeepFloyd IF: open-source text-to-image model

#118

Earlier quoted context omitted.

Midjourney always look very aesthetic pleasing, I guess because of their RLHF tuning with Discord data... But it doesn't really follow prompts as well as Dall-e for example. But in the end, people want pretty pictures. So is a complicated situation.

The tweet they shared is from February and uses an outdated version of MJ, this is what I got from V5: https://i.imgur.com/0uxtZDe.png Midjourney does much better overall . Composition is neat, but MJ is so incredibly far ahead in terms of quality of output, it honestly doesn't matter if you have to go and do composition manually (and with new AI based tools, that's easier than ever too. Do a bad cut and paste job th…

The thing is... IF is currently just a base model, it will need serious fine-tuning before it will produce aesthetically pleasing images (like MJ certainly does).

It's interesting to see what IF can do in terms of composition, text rendering etc, it's very promising if aesthetically pleasing images can be achieved via fine-tuning (the same happened with SD... current publicly fine-tuned models can achieve much higher levels of quality and cohesion than the base models, here's the prompt in an SD2.1 based model: https://imgur.com/a/ELGMSmV ).

Of course fine-tuning IF is likely more challenging, as both the two first stages and the 4x SD upscaler might need to be fine-tuned...

Re: DeepFloyd IF: open-source text-to-image model

#119

Earlier quoted context omitted.

But it didn't work? In yours there is no "Nexus", no smiling, no frowning, and man on right doesn't look Asian? Compared with image in Tweet, MJ failed at this task.

In what way does the man on the right not look like he could be from the absolutely enormous continent known as Asia?

A human interpreting the prompt would see "asian" as being in contrast to "indian" in the language of the prompt... Not a level of comprehension that can be expected of current models but maybe in a few years (months?).

Re: DeepFloyd IF: open-source text-to-image model

#120

Example of how much better it can do compared to midjourney, on a complex prompt: https://twitter.com/eb_french/status/1623823175170805760 It is able to put people on the left/right and put the correct t-shirts and facial expressions on each one. This is compared to mj which just mixes together a soup of every word you use and plops it out into the image. Huge MJ fan of course, it's amazing, but having compositional…

Spatial composition can be done easily, if you stop bothering with pure text-to-image (SD has several tricks and UIs to place objects precisely, they are all janky but they do work, that's practically photobashing). Attribute separation is also easily done with tricks like token bucketing, so your Indian guy will look Indian, and your East Asian guy will look East Asian. All of that is easy if you abandon the ambiguous natural language and use higher-order guidance.

What's really required is semantic composition. Making subjects meaningfully and predictably interact, or combining them together. And also the coherence of the overall stitched picture, so you don't end up with several different perspective planes.

Post reply on HN