Live data from Hacker News

Comparing Adobe Firefly, Dalle-2, and OpenJourney

blog.usmanity.com

11–20 of 139 posts

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#11

not sure this is a good comparison. midjourney likes much shorter prompts, and honestly they're all absolutely terrible for anything that isn't 'photo' based. E.g. ask it to generate a word bubble of the most common programming languages and it will fail every time, no matter what you try. I love it for photo stuff, but for photoshop you'd expect it to be able to do other things as well.

Curious, midjourney does great art and cartoon/comic styles too. Not just realistic images.

Most image AI tools are terrible with words.

I am curious, what images did you try generating with midjourney?

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#12
post #3

Midjurney is still so far ahead it's no competition. Did a lot of testing today and firefly generated so much errors with fingers and stuff, not seen that since the original stability release. Anyone know if the web firefly and the Photoshop version is the same model?

Not with typography though, haha. It can't spell. I had to draw the letters myself

None of these can do text well. There's a model that does do text and composition well, but the name escapes me. And the general quality is much lower overall, so it's a pretty heavy tradeoff.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#13
I had done a similar comparison a couple months back but used Lexica instead of DALL-E.

Seems clear to me that Midjourney has by far the best "vibes" understanding. Most models get the items right but not the lighting. Firefly seems focused on realism which makes sense for a photography audience.

https://twitter.com/fanahova/status/1639325389955952640?s=46...

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#14

not sure this is a good comparison. midjourney likes much shorter prompts, and honestly they're all absolutely terrible for anything that isn't 'photo' based. E.g. ask it to generate a word bubble of the most common programming languages and it will fail every time, no matter what you try. I love it for photo stuff, but for photoshop you'd expect it to be able to do other things as well.

That’s not a fair comparison, as Midjourney is outstanding at a wide range of styles beyond photography.

Generating a “word bubble” is going to look terrible in every major diffusion model. Cohesive words and writing in image models is still highly specialised.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#15
For those curious, I tried the same prompts with Kandinsky 2.1 [0]. In my experience it kind of blends the conceptual understanding of DALL-E with the higher quality image generation of Stable Diffusion. Like Midjourney though it kind of injects it's own style and allows you to get "satisfying" results from short prompts.

The flaw with these comparisons is that you really shouldn't use the same prompt with different generators. If you want to get best results you do have to play with the prompts and do a bunch of iteration to kind of explore the latent space and find what you're looking for. The first super long prompt looks like it's tuned for stable diffusion for instance. Different generators also have different syntax (e.g. with stable diffusion you can surround a phrase with parens to give it extra emphasis).

[0]: https://iterate.world/s/clj4n19u20000jv08iqygiaqw

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#16
For reference, here's what you can get with a properly tweaked Stable Diffusion, all running locally on my PC. Can be set up on almost any PC with a mid range GPU in a few minutes if you know what you're doing. I didn't do any cherry picking; this is the first thing it generated. 4 images per prompt.

1st prompt: https://i.postimg.cc/T3nZ9bQy/1st.png

2nd prompt: https://i.postimg.cc/XNFm3dSs/2nd.png

3rd prompt: https://i.postimg.cc/c1bCyqWR/3rd.png

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#17
post #10

Amazing how quickly Dalle-2 went from among the best image transformers to among the worst.

The stagnation has been very curious. They are part of a large & generally competent org, which otherwise has remained far ahead of the competition, like GPT-4. Except... for DALL-E 2, where it did not just stagnate for over a year (on top of its bizarre blindspots like garbage anime generation), but actually seemed to get worse . They have an experimental model of some sort that some people have access to, but even…

Nobody is able to use Parti or eDiff. Compared to models you can use, the experimental Dall-e or Bing Image Creator is second only to midjourney in my experience.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#19

Earlier quoted context omitted.

Not with typography though, haha. It can't spell. I had to draw the letters myself

None of these can do text well. There's a model that does do text and composition well, but the name escapes me. And the general quality is much lower overall, so it's a pretty heavy tradeoff.

DeepFloyd?

https://github.com/deep-floyd/IF

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#20

Amazing how quickly Dalle-2 went from among the best image transformers to among the worst.

I don't know, what I saw in there (particularly with the haunted house) was a far broader POTENTIAL RANGE of outputs. I get that they were cheesier outputs, but it seems to me that those outputs were just as capable of coming from the other 'AIs'… if you let them.

It's like each of these has a hidden giant pile of negative prompts, or additional positive prompts, that greatly narrow down the range of output. There are contexts where the Dall-E 'spoopy haunted house ooooo!' imagery would be exactly right… like 'show me halloweeny stock art'.

That haunted house prompt didn't explicitly SAY 'oh, also make it look like it's a photo out of a movie and make it look fantastic'. But something in the more 'competitive' AIs knew to go for that. So if you wanted to go for the spoopy cheesey 'collective unconscious' imagery, would you have to force the more sophisticated AIs to go against their hidden requirements?

Mind you if you added 'halloween postcard from out of a cheesey old store' and suddenly the other ones were doing that vibe six times better, I'd immediately concede they were in fact that much smarter. I've seen that before, too, in different Stable Diffusion models. I'm just saying that the consistency of output in the 'smarter' ones can also represent a thumb on the scale.

They've got to compete by looking sophisticated, so the 'Greg Rutkowskification' effect will kick in: you show off by picking a flashy style to depict rather than going for something equally valid, but less commercial.

Post reply on HN