Live data from Hacker News

DeepFloyd IF: open-source text-to-image model

github.com

121–130 of 237 posts

Re: DeepFloyd IF: open-source text-to-image model

#121

Earlier quoted context omitted.

> as an experienced user you should know that's it's not one shot Being "not one shot" for most nontrivial prompts is a failure of current t2i models, its what they all strive for and its what DF supposedly does a lot better. And, while its possible to spin things pretty hard when people can't bang on it themselves, I think the indication is that it is, in fact, a major leap forward from the best current consumer-ava…

Are you associated with deepfloyd? It's not a major leap not even a small one because it's exactly like imagen. It's stability giving some Ukrainian refugees compute time to train "their" model for publicity. It's about the whom and not what as it should be. I "feel" nothing I am telling it how it is. Look at the afghan girl example again it. Close up portrait, same clothing, same comp, expressive eyes... and most im…

> Are you associated with deepfloyd?

No, I’m not affiliated with StabilityAI

> It’s not a major leap not even a small one because it’s exactly like imagen.

I would agree, if imagen was a “consumer-available t2i model”. What’s available is a research paper with demo images from Google. The model itself is locked up inside Google, notionally because they haven’t solved filtering issues with it.

> Look at the afghan girl example again it. Close up portrait, same clothing, same comp, expressive eyes…

You look at it again, literally none of those things are the same: its not the same clothing (the material and color of the head scarf is different, the headscarf is the only visible clothing in the DF image, whereas that is not the case in the famous image), the condition of the head scarf is different, the hair color is different, the hair style is different, the hair texture is different, the face shape is different, the individual facial features are different, the eye color is much more brown in the DF image, the facial expression is different, the DF image has lipstick and eyeshadow, the famous image has a dirty face and no makeup, the headscarf is worn differently in the two images, the background is different, the lighting is different, and the faces are framed differently.

The similarities are (1) its a close up portrait, (2) a general ethnic similarity, and (3) they are both wearing a red (though very different red) head scarf, (4) and they are both looking straight into the camera. (2)-(4) are explicitly prompted, (1) is strongly implied in the prompt addressing nothing that isn’t related to the face/head. This isn’t “overfitting on a copyright image” its getting what you prompt, with no other similarity to the existing image.

> You guys all want it to be something special and I get it,

I’m actually kind of annoyed, because I’ve been collecting tooling, checkpoints, and other support for, and spending quite a bit of time getting proficient in dealing with the quirks of, Stable Diffusion. But, that’s life.

> it’s neither a good architecture nor a good implementation.

I’d be interested in hearing your specific criticism of the architecture and implementation, but hopefully its more grounded in fact than your criticism of the one image...

Re: DeepFloyd IF: open-source text-to-image model

#122

Earlier quoted context omitted.

Can't wait to see how all of this is going to look like in ten years. I know we are all nitpicking right now but these results are totally mindblowing already.

The thing is all these things can go downhill as well. All these things are cool today just like Google was a decade ago.

Even though most of these we can do locally?

Re: DeepFloyd IF: open-source text-to-image model

#123
post #119

Earlier quoted context omitted.

In what way does the man on the right not look like he could be from the absolutely enormous continent known as Asia?

A human interpreting the prompt would see "asian" as being in contrast to "indian" in the language of the prompt... Not a level of comprehension that can be expected of current models but maybe in a few years (months?).

I'm human. I interpreted Asian as from a non-specific part of Asia. I realise though that "Asian" has a very specific meaning in the US, but it's only the US that does this. For the rest of the world Asian means someone from Asia.

Re: DeepFloyd IF: open-source text-to-image model

#124
post #60

Earlier quoted context omitted.

Yep. It's getting really exhausting seeing projects falsely advertising themselves as "open source". Either be FOSS or don't be; don't pretend to be while using some nonsense like the BSL or whatever adhocery is in play here.

In the README they even call it "Modified MIT", the modification being where they turned it from a very permissive license into a fully proprietary one. Very cool model though.

Well, the Apache Software Foundation were rightly getting annoyed at 'modified Apache' licences...

Re: DeepFloyd IF: open-source text-to-image model

#125

Earlier quoted context omitted.

But it didn't work? In yours there is no "Nexus", no smiling, no frowning, and man on right doesn't look Asian? Compared with image in Tweet, MJ failed at this task.

In what way does the man on the right not look like he could be from the absolutely enormous continent known as Asia?

Thing is, it's possible to just run experiments against your claim that "actually the bot was doing a good job"

https://twitter.com/eb_french/status/1651584746089218049

it wasn't.

Again, I don't know why everyone is so defensive. I love MJ. There's nothing wrong with admitting that other models might do certain things better. We all can use any model we want.

Re: DeepFloyd IF: open-source text-to-image model

#127
post #108
post #94

Earlier quoted context omitted.

Where does it prohibit commercial use? I’m not seeing that in the license.

The model license, 1(a) and 2(a)(i).

Ah I guess I read it right then I reread it wrong, my bad and thanks for the pointer! That's a shame but hopefully it's released in a more open way in the future. My interest is in building good collaborative interfaces (and games!) on top of these things.

Re: DeepFloyd IF: open-source text-to-image model

#128
post #119

Earlier quoted context omitted.

A human interpreting the prompt would see "asian" as being in contrast to "indian" in the language of the prompt... Not a level of comprehension that can be expected of current models but maybe in a few years (months?).

I'm human. I interpreted Asian as from a non-specific part of Asia. I realise though that "Asian" has a very specific meaning in the US, but it's only the US that does this. For the rest of the world Asian means someone from Asia.

That's kind of what I meant... If a prompt specifies one "Indian", and one "Asian", that implies that the writer of the prompt doesn't think of Indian as Asian so probably from the US background.

Re: DeepFloyd IF: open-source text-to-image model

#129

Earlier quoted context omitted.

In what way does the man on the right not look like he could be from the absolutely enormous continent known as Asia?

Thing is, it's possible to just run experiments against your claim that "actually the bot was doing a good job" https://twitter.com/eb_french/status/1651584746089218049 it wasn't. Again, I don't know why everyone is so defensive. I love MJ. There's nothing wrong with admitting that other models might do certain things better. We all can use any model we want.

There's a certain irony in tweeting ad hominem attacks then claiming people are being defensive...

Re: DeepFloyd IF: open-source text-to-image model

#130

Earlier quoted context omitted.

Thing is, it's possible to just run experiments against your claim that "actually the bot was doing a good job" https://twitter.com/eb_french/status/1651584746089218049 it wasn't. Again, I don't know why everyone is so defensive. I love MJ. There's nothing wrong with admitting that other models might do certain things better. We all can use any model we want.

There's a certain irony in tweeting ad hominem attacks then claiming people are being defensive...

Yeah perhaps I was not good at judging tone. To me it's a matter of fact thing that MJ isn't good at this. They'd admit it, it's not a big deal and I'm a fan.

It's not ad hominem at all... mj isn't as good at certain types of composition as others. I don't get why people are pretending that isn't the case. I want everyone to have great models and IF is part of that progress. Perhaps calling it "word soup" was offensive? This isn't your religion, though, it's just a model. Listening in on the MJ office hours they're the farthest thing you can be from dogmatic or arrogant. They want to improve as we all do. I personally am just really inspired that everyone can advance together!

Also see downthread - the first 32 images I generated attempting to reproduce the claim that "actually MJ can do this" all failed. The person who challenged me then ignored it. This isn't really up for debate until someone sends a seed where mj can do the cube + sphere thing well.

Post reply on HN