Live data from Hacker News

DeepFloyd IF: open-source text-to-image model

github.com

91–100 of 237 posts

Re: DeepFloyd IF: open-source text-to-image model

#91

Earlier quoted context omitted.

> it means the model licensor has a claim against you Right, but to what end? The only reason the licensor should care one way or another is the licensor being held liable for what folks do with the software, in which case... > it also means, and this is probably more important, that in some juridictions, the model licensor has a better defense against liability for contributory infringement if the licensee infringes…

> Do hardware stores need to demand "thou shalt not use this tool to kill people" to their customers to avoid liability for axe murders under such jurisdictions? Generally, not, because vicarious liability for battery and wrongful death doesn’t work like, e.g., contributory copyright infringement. > I'm pretty sure the standard warranty disclaimer in your average FOSS license already covers this No, warranty disclaim…

> Generally, not, because vicarious liability for battery and wrongful death doesn’t work like, e.g., contributory copyright infringement.

Judging by youtube-dl, it seems like it does work that way, at least in my jurisdiction; I guess we'll see if the RIAA doubles down on trying to wipe it from the face of the Earth, but considering there hasn't been much noise, I wouldn't count on it. Also, to my adjacent point, I highly doubt the RIAA would've refrained from attempting to take down youtube-dl even if youtube-dl's license prohibited its users from circumventing DRM with it.

Re: DeepFloyd IF: open-source text-to-image model

#93
post #90

Earlier quoted context omitted.

You think? Automatic1111 is still on pytorch 1.7 and SD1.5

Stable Diffusion 2.x has been supported for a while.

Yup, lots of misinformation in this thread from those who are not in-the-know.

Automatic1111 is the defacto main UI for these kind of models. It will be supported there, quite quickly.

Re: DeepFloyd IF: open-source text-to-image model

#94
post #41
post #30

Earlier quoted context omitted.

> model license prohibits commercial use I thought that at first, but I think it only prohibits commercial use that breaks regional copyright or privacy laws.

It prohibits both commercial use, whether or not you break regional laws; and it prohibits breaking certain laws. As another user said, encoding the law into a licence is pointless but makes it non-free. There are also problematic restrictions on your ability to modify the software under clause 2(c). And nor do you have the right to sublicence, it's not clear to me what rights somebody has if you give them a copy.

Where does it prohibit commercial use? I’m not seeing that in the license.

Re: DeepFloyd IF: open-source text-to-image model

#95

Earlier quoted context omitted.

Its not pointless, it means the model licensor has a claim against you, as well as whoever would for violating the referenced laws; it also means, and this is probably more important, that in some juridictions, the model licensor has a better defense against liability for contributory infringement if the licensee infringes. EDIT: That said, it’s unambiguously not open source.

> it means the model licensor has a claim against you Right, but to what end? The only reason the licensor should care one way or another is the licensor being held liable for what folks do with the software, in which case... > it also means, and this is probably more important, that in some juridictions, the model licensor has a better defense against liability for contributory infringement if the licensee infringes…

> The only reason the licensor should care one way or another is the licensor being held liable for what folks do with the software, in which case...

There are current open cases of people claiming “harm” for misinformation spouted by ChatGPT where it just makes up facts to satisfy user prompts.

There are current open cases where people are claiming “copyright violation” due to diffusion models satisfying user prompts.

AFAICT none of these cases are against the users who are prompting the models.

Re: DeepFloyd IF: open-source text-to-image model

#97

Example of how much better it can do compared to midjourney, on a complex prompt: https://twitter.com/eb_french/status/1623823175170805760 It is able to put people on the left/right and put the correct t-shirts and facial expressions on each one. This is compared to mj which just mixes together a soup of every word you use and plops it out into the image. Huge MJ fan of course, it's amazing, but having compositional…

Can't wait to see how all of this is going to look like in ten years. I know we are all nitpicking right now but these results are totally mindblowing already.

Re: DeepFloyd IF: open-source text-to-image model

#98
post #79

Earlier quoted context omitted.

Interesting, I do not get the results you do. What additional parameters are you using? Here is a link to some of my tests, with all default settings, some in v5 some in v4. https://twitter.com/eb_french/status/1651370091869786112 0/16 images have a red cube on a green sphere.

none and as an experienced user you should know that's it's not one shot and most of the time not even few shot... You can't compare cherry picked press images with few shots of a 5 second prompt. I don't know why you want to hype something up if you can't really compare it. It seems extremly attention grifting. Just look at their cherry picks in this discord... https://discord.com/invite/pxewcvSvNx . It's overfitted…

> as an experienced user you should know that's it's not one shot

Being "not one shot" for most nontrivial prompts is a failure of current t2i models, its what they all strive for and its what DF supposedly does a lot better. And, while its possible to spin things pretty hard when people can't bang on it themselves, I think the indication is that it is, in fact, a major leap forward from the best current consumer-available t2i models (it looks pretty comparable to Google Imagen – a little bit worse benchmark scores – which is unsurprising since it seems to be an implementation of exactly the architecture described in Google's Imagen paper.

> It’s overfitted on images with copyright (afghan girl)

It’s…not, though. Sure, the picture with a prompt which is suggestive of that (down to even specifying the same film type) gives off a vibe that completely feels, if you haven’t recently looked at the famous picture but are familiar with it, like a “cleaned up” version of that picture, so you might intuitively feel its from overfitting, that it is basically reproducing the original image with slight variations.

Look at the two pictures side-by-side and there is basically nothing similar about them except exactly the things specified in the prompt, and pretty much every aspect of the way that those elements of the prompt is interpreted in the DF image is unlike the other image.

Re: DeepFloyd IF: open-source text-to-image model

#99
post #29

For anyone who doesn't know, DeepFloyd is a StableDiffusion style image model that more or less replaced CLIP with a full LLM (11b params). The result is that it is much better at responding to more complex prompts. In theory, it is also smarter at learning from its training data.

It isn't like Stable Diffusion, it's more like Google's Imagen model.

> It isn’t like Stable Diffusion, it’s more like Google’s Imagen model.

Yeah, it looks exactly (architecturally) like Imagen.

Google would be running circles around everyone in Generative AI (maybe OpenAI would still have a better core LLM, maybe, but portfolio-wise) if they simply had the ability to cross the gap between building technologies and writing up research papers on them and actually releasing products.

Re: DeepFloyd IF: open-source text-to-image model

#100

> Text > Hands good god it solves the two biggest meme issues with image models in one go. Will this be the new state of the art every other model is compared to?

There's fundamental tradeoffs, as there always will be when you're compressing things into an image model. So, this is going to have new different issues. Since it's similar to Imagen, it probably can't handle long complex prompts as well, since they developed Parti afterward. Here's my question: are there any image models where, if you prompt "1+1", you get an image showing "3"?

> So, this is going to have new different issues.

Well, yeah, its a bigger set of models (particular the language model) that takes more resources (both to train and for inference.) That’s the tradeoff.

> Here’s my question: are there any image models where, if you prompt “1+1”, you get an image showing “3”?

You want a t2i model that does arithmetic in the prompt, translates to it to “text displaying the number ”, but, also does the arithmetic wrong?

Yeah, I don’t think that combination of features is in any existing model or, really, in any of the datasets used for evaluation, or otherwise on anyone’s roadmap.

Post reply on HN