Live data from Hacker News

DeepFloyd IF: open-source text-to-image model

github.com

221–230 of 237 posts

Re: DeepFloyd IF: open-source text-to-image model

#221
post #167

Earlier quoted context omitted.

It's much better, but it's not perfect. Here's what I got for: > a photograph of raccoon in the woods holding a sign that says "I will eat your trash" https://twitter.com/simonw/status/1651994059781832704

This is not a problem. The sign was clearly made by the raccoon.

Agreed - to properly test it, we should try:

A photograph of an English professor in the woods holding a sign that says "I will eat your trash"

Re: DeepFloyd IF: open-source text-to-image model

#222

Earlier quoted context omitted.

What does this mean? Isn't the quality of a model determined by how easy it is to get a good picture?

Not necessarily. IMO a good model needs to follow your prompt well, and that was my problem with Stable Diffusion. I've been trying to get a good portrait picture with "neon lights" on Stable Diffusion and it is almost impossible. Meanwhile with the new Dall-e, that was possible. The picture specially with SDXL is good, but it doesn't really have neon lights... I tried now similar prompt on deepfloyd and managed to g…

would be interesting if you could used deepfloyd first for image composition, then apply stable diffusion after for purely stylistic modifications

Re: DeepFloyd IF: open-source text-to-image model

#223

Earlier quoted context omitted.

Not necessarily. IMO a good model needs to follow your prompt well, and that was my problem with Stable Diffusion. I've been trying to get a good portrait picture with "neon lights" on Stable Diffusion and it is almost impossible. Meanwhile with the new Dall-e, that was possible. The picture specially with SDXL is good, but it doesn't really have neon lights... I tried now similar prompt on deepfloyd and managed to g…

would be interesting if you could used deepfloyd first for image composition, then apply stable diffusion after for purely stylistic modifications

Definitely possible :) I've been doing this with new Dall-e + img2img with Stable Diffusion.

Explaining: I created a model of me, and wanted to create some good realistic portrait pictures. First I tried to create a model of me using some of the custom models already exist and the result was bad.

Then I tried SD 1.5/2.1... It was better, but couldn't really get some of the prompts make real...

Then I tried new Dall-e, saved, and inserted my face with img2img on SD and it worked much better!

Re: DeepFloyd IF: open-source text-to-image model

#224

Earlier quoted context omitted.

Does the increased memory footprint mean it can't be run on a normal desktop like SD?

The VRAM requirements are higher (14GB) so lots of things that can do SD won’t do this with thr existing toolchain. But some of that is “aftermarket” SD optimization, and this maybe could see some of that, too. But there are consumer cards with 14GB+ VRAM, so its not, even before optimization, out of reach of consumer hardware.

Damn. That's basically just the 4090 and 4080.

Re: DeepFloyd IF: open-source text-to-image model

#225

Any web based front ends yet? I put together a system that runs a variety of web based open source AI image generation and editing tools on Vultr GPU instances. It spins up instances on demand, mounts an NFS filesystem with local caching and a COW layer, spawns the services, proxies the requests, and then spins down idle instances when I'm done. Would love to add this, suppose I could whip something up if none exists…

What's your app / service called?!

Not publicly released. I'd be happy to show you if you send me an email. See my profile.

Re: DeepFloyd IF: open-source text-to-image model

#226

Earlier quoted context omitted.

The thing is all these things can go downhill as well. All these things are cool today just like Google was a decade ago.

Even though most of these we can do locally?

He's probably referring to the _business_ of search. The B word taints a lot of things.

Re: DeepFloyd IF: open-source text-to-image model

#227

Earlier quoted context omitted.

You can't use this to make logos for any commercial product, and it's not safe to use it for hobby projects either, based on the current model license.

You also cannot just ingest people's media to train a for-profit AI image generation service, and well, here we are. PS: not an AI apologist, just pointing the irony. Feels like those fan sonic characters "original content do not steal".

Yes you can. This is nonsense.

https://www.gesetze-im-internet.de/urhg/__44b.html

Re: DeepFloyd IF: open-source text-to-image model

#228

Earlier quoted context omitted.

You can't use this to make logos for any commercial product, and it's not safe to use it for hobby projects either, based on the current model license.

> You can't use this to make logos for any commercial product Yeah good luck figuring out that a particular logo was generated with this particular model. And if someone does good luck doing anything about it. With this amount of fear one wouldn't dare to cross a road without three layers of bubble wrap, plus written authorisation from a lawyer plus a feasibility study from a traffic engineer.

You've never gone through an acquisition or due diligence, have you?
Post reply on HN