Live data from Hacker News

Comparing Adobe Firefly, Dalle-2, and OpenJourney

blog.usmanity.com

61–70 of 139 posts

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#61

For reference, here's what you can get with a properly tweaked Stable Diffusion, all running locally on my PC. Can be set up on almost any PC with a mid range GPU in a few minutes if you know what you're doing. I didn't do any cherry picking; this is the first thing it generated. 4 images per prompt. 1st prompt: https://i.postimg.cc/T3nZ9bQy/1st.png 2nd prompt: https://i.postimg.cc/XNFm3dSs/2nd.png 3rd prompt: https:…

I am sure you're right, but "if you know what you're doing" does a lot of heavy lifting here.

We could just as easily say "hosting your own email can be set up in a few minutes if you know what you're doing". I could do that, but I couldn't get local SD to generate comparable images if my life depended on it.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#62
post #39

Earlier quoted context omitted.

Can you elaborate on “properly tweaked”? When I use one of the Stable Diffusion and AUTOMATIC1111 templates on runpod.io, the results are absolutely worthless. This is using some of the popular prompts you can find on sites like prompthero that show amazing examples. It’s been serious expectation vs. reality disappointment for me and so I just pay the MidJourney or DALL-E fees.

You're not going to get even close to Midjourney or even Bing quality on SD without finetuning. It's that simple. When you do finetune, it will be restricted to that aesthetic and you won't get the same prompt understanding or adherence. For all the promise of control and customization SD boasts, Midjourney beats it hands down in sheer quality. There's a reason like 99% of ai art comic creators stick to Midjourney de…

I feel like people shouldn't talk in definitives if their message is just going to demonstrate they have no idea what they're talking about.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#63
post #39

For reference, here's what you can get with a properly tweaked Stable Diffusion, all running locally on my PC. Can be set up on almost any PC with a mid range GPU in a few minutes if you know what you're doing. I didn't do any cherry picking; this is the first thing it generated. 4 images per prompt. 1st prompt: https://i.postimg.cc/T3nZ9bQy/1st.png 2nd prompt: https://i.postimg.cc/XNFm3dSs/2nd.png 3rd prompt: https:…

Can you elaborate on “properly tweaked”? When I use one of the Stable Diffusion and AUTOMATIC1111 templates on runpod.io, the results are absolutely worthless. This is using some of the popular prompts you can find on sites like prompthero that show amazing examples. It’s been serious expectation vs. reality disappointment for me and so I just pay the MidJourney or DALL-E fees.

> Can you elaborate on “properly tweaked”?

In a nutshell:

1. Use a good checkpoint. Vanilla stable diffusion is relatively bad. There are plenty of good ones on civitai. Here's mine: https://civitai.com/models/94176

2. Use a good negative prompt with good textual inversions. (e.g. "ng_deepnegative_v1_75t", "verybadimagenegative_v1.3", etc.; you can download those from civitai too) Even if you have a good checkpoint this is essential to get good results.

3. Use a better sampling method instead of the default one. (e.g. I like to use "DPM++ SDE Karras")

There are more tricks to get even better output (e.g. controlnet is amazing), but these are the basics.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#64

Earlier quoted context omitted.

You're not going to get even close to Midjourney or even Bing quality on SD without finetuning. It's that simple. When you do finetune, it will be restricted to that aesthetic and you won't get the same prompt understanding or adherence. For all the promise of control and customization SD boasts, Midjourney beats it hands down in sheer quality. There's a reason like 99% of ai art comic creators stick to Midjourney de…

You load a model and have 6 sliders instead of one… it’s not exactly “fine tuning”. If you want the power, it’s there. But nearly bone stock SD in auto1111 is going to get to any of these examples easily. Show me the civitai equivalent for MJ or Dalle2. It doesn’t exist.

>You load a model and have 6 sliders instead of one… it’s not exactly “fine tuning”.

Ok...? Read what i wrote carefully. Your 6 sliders won't produce better images than midjourney for your prompt on the base SD model.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#65

I like how simple Firefly’s images are, like something you’d want to work with in Photoshop. Dalle-2 looks terrible. Midjourney is still my favorite.

As someone who has spent hours playing with it in Photoshop (Beta) Firefly is actually pretty damned cool!

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#66
post #59

Earlier quoted context omitted.

Are you using txt2img with the vanilla model? SD's actual value is in the large array of higher-order input methods and tooling; as a tradeoff, it requires more knowledge. Similarly to 3D CGI, it's a highly technical area. You don't just enter the prompt with it. You can finetune it on your own material, or choose one of the hundreds of public finetuned models. You can guide it in a precise manner with a sketch or by…

Is there a coherent resource (not a scattered 'just google it' series of guides from all over the place) that encapsulates some of the concepts and workflows you're describing? What would be the best learning site/resource for arriving at understanding how to integrate and manipulate SD with precision like that? Thanks

> What would be the best learning site/resource for arriving at understanding how to integrate and manipulate SD with precision like that?

Honestly? Probably YouTube tutorials.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#67

Earlier quoted context omitted.

You're not going to get even close to Midjourney or even Bing quality on SD without finetuning. It's that simple. When you do finetune, it will be restricted to that aesthetic and you won't get the same prompt understanding or adherence. For all the promise of control and customization SD boasts, Midjourney beats it hands down in sheer quality. There's a reason like 99% of ai art comic creators stick to Midjourney de…

I feel like people shouldn't talk in definitives if their message is just going to demonstrate they have no idea what they're talking about.

I know what i'm talking about lol. I tuned a custom SD model that's downloaded thousands of times a month. I'm speaking from experience more than anything. Don't know why some SD users get so defensive.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#68
post #39

Earlier quoted context omitted.

Can you elaborate on “properly tweaked”? When I use one of the Stable Diffusion and AUTOMATIC1111 templates on runpod.io, the results are absolutely worthless. This is using some of the popular prompts you can find on sites like prompthero that show amazing examples. It’s been serious expectation vs. reality disappointment for me and so I just pay the MidJourney or DALL-E fees.

You're not going to get even close to Midjourney or even Bing quality on SD without finetuning. It's that simple. When you do finetune, it will be restricted to that aesthetic and you won't get the same prompt understanding or adherence. For all the promise of control and customization SD boasts, Midjourney beats it hands down in sheer quality. There's a reason like 99% of ai art comic creators stick to Midjourney de…

Yet you are posting this in a thread where GP provided actual examples of the opposite. Look for another comment above/below, there are MJ-generated samples which are comparable but also less coherent than the result from a much smaller SD model. And in case of MJ hallucinations cannot be fixed. MJ is good but it isn't magic, it just provides quick results with little experience required; prompt understanding is still poor, and will stay poor until it's paired with a good LLM.

Neither of the existing models gives actually passable production-quality results, be it MJ or SD or whatever else. It will be quite some time until they get out of the uncanny valley.

> There's a reason like 99% of ai art comic creators stick to Midjourney

They aren't. MJ is mostly used by people without experience, think a journalist who needs a picture for an article. Which is great and it's what makes them good money.

As a matter of fact (I work with artists), for all the surface-visible hate AI art gets in the artist community, many actual artists are using it more and more to automate certain mundane parts of their job to save time, and this is not MJ or Dall-E.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#69

Earlier quoted context omitted.

You're not going to get even close to Midjourney or even Bing quality on SD without finetuning. It's that simple. When you do finetune, it will be restricted to that aesthetic and you won't get the same prompt understanding or adherence. For all the promise of control and customization SD boasts, Midjourney beats it hands down in sheer quality. There's a reason like 99% of ai art comic creators stick to Midjourney de…

Yet you are posting this in a thread where GP provided actual examples of the opposite. Look for another comment above/below, there are MJ-generated samples which are comparable but also less coherent than the result from a much smaller SD model. And in case of MJ hallucinations cannot be fixed. MJ is good but it isn't magic, it just provides quick results with little experience required; prompt understanding is stil…

>Yet you are posting this in a thread where GP provided actual examples of the opposite.

Opposite of what ? OP posts results from a tuned model.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#70
post #10

Amazing how quickly Dalle-2 went from among the best image transformers to among the worst.

The stagnation has been very curious. They are part of a large & generally competent org, which otherwise has remained far ahead of the competition, like GPT-4. Except... for DALL-E 2, where it did not just stagnate for over a year (on top of its bizarre blindspots like garbage anime generation), but actually seemed to get worse . They have an experimental model of some sort that some people have access to, but even…

I suspect that they consider txt2img to be more of a curiosity now. Sure, it's transformative; it's going to upend whole markets (and make some people a lot of money in the process) - however, it's just producing images. Contrast with LLMs, which have already proven to be generally applicable in great many domains, and that if you squint, are probably capturing the basic mechanisms of thinking. OpenAI lost the lead in txt2img, but GPT-4 is still way ahead of every other LLM. It makes sense for them to focus pretty much 100% on that.
Post reply on HN