Live data from Hacker News

DeepFloyd IF: open-source text-to-image model

github.com

81–90 of 237 posts

Re: DeepFloyd IF: open-source text-to-image model

#81
post #59

Has anyone tried the Scott Alexander AI bet prompts? 1. A stained glass picture of a woman in a library with a raven on her shoulder with a key in its mouth 2. An oil painting of a man in a factory looking at a cat wearing a top hat 3. A digital art picture of a child riding a llama with a bell on its tail through a desert 4. A 3D render of an astronaut in space holding a fox wearing lipstick 5. Pixel art of a farmer…

where are these prompts from?

Scott Alexander made a bet with those prompts here: https://astralcodexten.substack.com/p/a-guide-to-asking-robo...

And followed up with this article when he won the bet: https://astralcodexten.substack.com/p/i-won-my-three-year-ai...

Re: DeepFloyd IF: open-source text-to-image model

#83

There's a discord with tons of sample images, where we've been waiting patiently for the release, coming SOON, for 3 months now. https://discord.gg/pxewcvSvNx

What these AI companies need are some good old-fashioned leakers. We should be seeing these models show up on sketchy pirate sites, complete with garish 80s-style cracking screens crediting various '1337 haX0rs with witty pseudonyms.

Re: DeepFloyd IF: open-source text-to-image model

#84

> Text > Hands good god it solves the two biggest meme issues with image models in one go. Will this be the new state of the art every other model is compared to?

There's fundamental tradeoffs, as there always will be when you're compressing things into an image model.

So, this is going to have new different issues. Since it's similar to Imagen, it probably can't handle long complex prompts as well, since they developed Parti afterward.

Here's my question: are there any image models where, if you prompt "1+1", you get an image showing "3"?

Re: DeepFloyd IF: open-source text-to-image model

#85
post #51

Any web based front ends yet? I put together a system that runs a variety of web based open source AI image generation and editing tools on Vultr GPU instances. It spins up instances on demand, mounts an NFS filesystem with local caching and a COW layer, spawns the services, proxies the requests, and then spins down idle instances when I'm done. Would love to add this, suppose I could whip something up if none exists…

It'll probably be in the Auto1111 WebUI within a week.

You think? Automatic1111 is still on pytorch 1.7 and SD1.5

Re: DeepFloyd IF: open-source text-to-image model

#86

Example of how much better it can do compared to midjourney, on a complex prompt: https://twitter.com/eb_french/status/1623823175170805760 It is able to put people on the left/right and put the correct t-shirts and facial expressions on each one. This is compared to mj which just mixes together a soup of every word you use and plops it out into the image. Huge MJ fan of course, it's amazing, but having compositional…

Midjourney always look very aesthetic pleasing, I guess because of their RLHF tuning with Discord data... But it doesn't really follow prompts as well as Dall-e for example.

But in the end, people want pretty pictures. So is a complicated situation.

Re: DeepFloyd IF: open-source text-to-image model

#87

Example of how much better it can do compared to midjourney, on a complex prompt: https://twitter.com/eb_french/status/1623823175170805760 It is able to put people on the left/right and put the correct t-shirts and facial expressions on each one. This is compared to mj which just mixes together a soup of every word you use and plops it out into the image. Huge MJ fan of course, it's amazing, but having compositional…

Midjourney always look very aesthetic pleasing, I guess because of their RLHF tuning with Discord data... But it doesn't really follow prompts as well as Dall-e for example. But in the end, people want pretty pictures. So is a complicated situation.

The tweet they shared is from February and uses an outdated version of MJ, this is what I got from V5: https://i.imgur.com/0uxtZDe.png

Midjourney does much better overall. Composition is neat, but MJ is so incredibly far ahead in terms of quality of output, it honestly doesn't matter if you have to go and do composition manually (and with new AI based tools, that's easier than ever too. Do a bad cut and paste job then infill your way back to a coherent image)

Re: DeepFloyd IF: open-source text-to-image model

#88

Earlier quoted context omitted.

Its not pointless, it means the model licensor has a claim against you, as well as whoever would for violating the referenced laws; it also means, and this is probably more important, that in some juridictions, the model licensor has a better defense against liability for contributory infringement if the licensee infringes. EDIT: That said, it’s unambiguously not open source.

> it means the model licensor has a claim against you Right, but to what end? The only reason the licensor should care one way or another is the licensor being held liable for what folks do with the software, in which case... > it also means, and this is probably more important, that in some juridictions, the model licensor has a better defense against liability for contributory infringement if the licensee infringes…

> Do hardware stores need to demand "thou shalt not use this tool to kill people" to their customers to avoid liability for axe murders under such jurisdictions?

Generally, not, because vicarious liability for battery and wrongful death doesn’t work like, e.g., contributory copyright infringement.

> I'm pretty sure the standard warranty disclaimer in your average FOSS license already covers this

No, warranty disclaimers don’t cover this, because (1) its not a warranty issue, and (2) disclaimers, if they have legal effect at all, effect liability the disclaiming party would otherwise have to the party accepting the disclaimer, not liability the disclaiming party would have to third parties.

Re: DeepFloyd IF: open-source text-to-image model

#89

There's a discord with tons of sample images, where we've been waiting patiently for the release, coming SOON, for 3 months now. https://discord.gg/pxewcvSvNx

What these AI companies need are some good old-fashioned leakers. We should be seeing these models show up on sketchy pirate sites, complete with garish 80s-style cracking screens crediting various '1337 haX0rs with witty pseudonyms.

Well, NovelAI was hacked and had their image generation model leaked last year.
Post reply on HN