Live data from Hacker News

Comparing Adobe Firefly, Dalle-2, and OpenJourney

blog.usmanity.com

111–120 of 139 posts

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#111
post #10

Earlier quoted context omitted.

The stagnation has been very curious. They are part of a large & generally competent org, which otherwise has remained far ahead of the competition, like GPT-4. Except... for DALL-E 2, where it did not just stagnate for over a year (on top of its bizarre blindspots like garbage anime generation), but actually seemed to get worse . They have an experimental model of some sort that some people have access to, but even…

I suspect that they consider txt2img to be more of a curiosity now. Sure, it's transformative; it's going to upend whole markets (and make some people a lot of money in the process) - however, it's just producing images. Contrast with LLMs, which have already proven to be generally applicable in great many domains, and that if you squint, are probably capturing the basic mechanisms of thinking . OpenAI lost the lead…

I find it curious because (a) if they don't care about text2image, why launch it as a service to begin with? (b) if they don't care now, why keep it up and let it keep consuming resources, human & GPU? (c) if they do still care, because as other models & services have demonstrated there's a ton of interest in text2image, why not invest the relatively minor amount of resources it would take to keep it competitive (look how few people work at Midjourney, or are authors on imagegen papers)? It may have cost >$100m to make GPT-4, but making a decent imagegen model costs a lot less than that! (Even now, you could probably create a SOTA model for But launching it and then just letting it stagnate indefinitely and get worse every day compared to its increasingly popular competitors seems like the worst of all worlds, and I can't see what is the OA strategy there.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#112
post #10

Earlier quoted context omitted.

The stagnation has been very curious. They are part of a large & generally competent org, which otherwise has remained far ahead of the competition, like GPT-4. Except... for DALL-E 2, where it did not just stagnate for over a year (on top of its bizarre blindspots like garbage anime generation), but actually seemed to get worse . They have an experimental model of some sort that some people have access to, but even…

Nobody is able to use Parti or eDiff. Compared to models you can use, the experimental Dall-e or Bing Image Creator is second only to midjourney in my experience.

Parti/eDiff show that it is relatively easy to do much better than the experimental model which presumably represents their best effort, never mind the hot garbage you see in OP from DALL-E 2. And it's not a calculated degree of low-quality enabled by those models being unreleased and having no competition, because competition like Stable Diffusion or Midjourney are beating the heck out of DALL-E 2 in popular usage.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#113
post #87

Earlier quoted context omitted.

I'm using my own custom trained model. Here, I've uploaded it to civitai: https://civitai.com/models/94176 There are plenty of other good models too though.

Any tips or guides you followed on training your custom model? I've done a few LoRAs and TI but haven't gotten to my own models yet. Your results look great and I'd love a little insight into how you arrived there and what methods/tools you used.

I'm not an expert at this and there are probably better ways to do this/might not work for you/your mileage may vary, so please take this with a huge grain of salt, but roughly this worked for me:

1. Start with a good base model(s) from which to train from.

2. Have a lot of diverse images.

3. Ideally train for only one epoch. (Having a lot of images helps here.)

4. If you get bad results lower the learning rate and try again.

5. After training try to mix your finetuned model with the original one, in steps of 10%, generate X/Y plot of it, pick the best result.

6. Repeat this process as long as you're getting an improvement.

For training I mostly used scripts from here: https://github.com/bmaltais/kohya_ss

The main problem here is that essentially during inference you're using a bag of tricks to make the output better (e.g. good negative embeddings), but when training you don't. (And I'm not entirely sure how you'd actually integrate those into the training process; might be possible, but I didn't want to spend too much time on it.) So your fine tuning as-is might improve the output of the model when no tricks are used, but it can also regress it when the tricks are used. Which I why I did the "mix and pick the best one" step.

But, again, I'm not an expert at this and just did this for fun. Ultimately there might be better ways to do it.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#114
post #92

Earlier quoted context omitted.

> Can you elaborate on “properly tweaked”? In a nutshell: 1. Use a good checkpoint. Vanilla stable diffusion is relatively bad. There are plenty of good ones on civitai. Here's mine: https://civitai.com/models/94176 2. Use a good negative prompt with good textual inversions. (e.g. "ng_deepnegative_v1_75t", "verybadimagenegative_v1.3", etc.; you can download those from civitai too) Even if you have a good checkpoint t…

What kind of(and how much) data did you use to train your checkpoint? I'd like to have a go at making one myself targeted towards single objects (be it car,spaceship, dinner plate, apple, octopus, etc). Most checkpoints are very heavily leaning towards people and portraits.

I’m not the OP but I’ve made some of my daughter, wife, dog, niece, etc.

People generally suggest 30+ images. I’ve found - at least with people - the more the better. My wife’s model is trained on ~80 images of her.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#115
post #28
post #3

Midjurney is still so far ahead it's no competition. Did a lot of testing today and firefly generated so much errors with fingers and stuff, not seen that since the original stability release. Anyone know if the web firefly and the Photoshop version is the same model?

I'm presuming you're not including Stable Diffusion when you say this; the fact that SD and its variants are defacto extremely "free and open source" presently put it way ahead of anything else, and are likely to do so for some time.

As far as I can tell anyone who’s creating images is using midjourney. This is likely the same “Linux is open so it’s way better” tell that to the trillion dollar companies that bet against that.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#116
post #111

Earlier quoted context omitted.

I suspect that they consider txt2img to be more of a curiosity now. Sure, it's transformative; it's going to upend whole markets (and make some people a lot of money in the process) - however, it's just producing images. Contrast with LLMs, which have already proven to be generally applicable in great many domains, and that if you squint, are probably capturing the basic mechanisms of thinking . OpenAI lost the lead…

I find it curious because (a) if they don't care about text2image, why launch it as a service to begin with? (b) if they don't care now, why keep it up and let it keep consuming resources, human & GPU? (c) if they do still care, because as other models & services have demonstrated there's a ton of interest in text2image, why not invest the relatively minor amount of resources it would take to keep it competitive (loo…

Maybe they keep it up just so that they have something in txt2img space? It may not be the best, or even good, but you don't know that until you try it, and until then, it just enhances the value of the OpenAI platform. E.g. if you're building something backed by OpenAI LLMs, and are thinking about future txt2img integration, the existence of Dall-E might stop you from "shopping around" txt2img services in advance.

The way I see it, they don't need txt2img at this moment - GPT-4 ensures they're the top #1 name both in the industry and in AI-related news stories. But it doesn't mean they won't come back to it. Couple observations:

- OpenAI isn't a "release early, release often" shop. They might be already working on something, but they'll release it only when it is a qualitative improvement over everyone else (or at least Dall-E).

- A bunch of hobbyists is doing all their work for free anyway. Stable Diffusion itself may not be SOTA, but the totality of hundreds of different fine-tunes on Civitai very much is. With all those models being shared in the open and relatively easy/cheap to recreate, it would make sense for OpenAI to just stand by and watch, and only invest resources once hobbyists hit a plateau.

- Looking at those Civitai models, it seems to me that OpenAI could beat txt2img SOTA easily, at any moment, by taking (or re-creating, depending on the license) the best five to ten SD derivatives, and put them behind GPT-4, or even GPT-3.5, fine-tuned to 1) choose the best SD derivative for user's prompt, and 2) transform user's prompt to set of parameters (positive & negative prompts, diffuser algo, numeric params) crafted with choice from 1) in mind. It's a black box. On the Internet, no one can tell you're an ensemble model.

- They could even be doing it as we speak - addition of function calls is aligned with this direction, fine-tuning for good prompt generation is mostly a txt2txt exercise, and again, hobbyists around the world are busy building a high-quality human-curated data set of {what I want}x{model + positive prompt + negative prompt + diffuser + other params} -> {is this any good?}. If I were them, I'd just mine this and not say anything.

- Overall, I think that in txt2img space, currently the hard part isn't the "img" part, but the "txt" part. OpenAI has a huge advantage here, and as long as its true, they're in position to instantly overtake everyone else in this space. That is, they have an "Ultimate attack" charged and ready, and are patiently waiting for a good moment to trigger it.

- Didn't they hint that GPT-4 successor will be multimodal? That could end up being their comeback to txt2img. And img2txt. And a bunch of other modalities.

EDIT: As if on cue, the very thing I was speculating about above is being discussed wrt. LLMs right now:

- https://news.ycombinator.com/item?id=36413296 - GPT-4 is 8 GPTs in a trench coat

- https://news.ycombinator.com/item?id=36413768 - 3-4 orders of magnitude efficiency (size vs effect) improvement in code generation, if your training data isn't garbage

And in both threads, people bring up older papers and discuss the merits of combining smaller specialized models into a more generic whole.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#117
post #28

Earlier quoted context omitted.

I'm presuming you're not including Stable Diffusion when you say this; the fact that SD and its variants are defacto extremely "free and open source" presently put it way ahead of anything else, and are likely to do so for some time.

As far as I can tell anyone who’s creating images is using midjourney. This is likely the same “Linux is open so it’s way better” tell that to the trillion dollar companies that bet against that.

To be honest most of the AI generated images I find online are generated by Stable Diffusion, the fact that you can't generate NSFW images with MJ makes also a big difference.

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#118

Earlier quoted context omitted.

> What would be the best learning site/resource for arriving at understanding how to integrate and manipulate SD with precision like that? Honestly? Probably YouTube tutorials.

Jaysus. I'm going to sound like an entitled whiny old guy shouting at clouds, but - what the hell; with all the knowledge being either locked and churned on Discord, or released in form of YouTube videos with no transcript and extremely low content density - how is anyone with a job supposed to keep up with this? Or is that a new form of gatekeeping - if you can't afford to burn a lot of time and attention as if in s…

The difference being that youtube videos can make more money for the author. Anyway, it's all open source, so feel free to make a wiki

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#119
post #79
post #78

Earlier quoted context omitted.

Would mid-journey be liable though? I mean you can create copyrighted material using photoshop too. (Even paint!). If I create a Mickey Mouse using photoshop would adobe be liable for it?

I don't think it really matters whether or not Midjourney themselves are liable, the output of their model being legally radioactive would break their business model either way. They make money by charging users for commercial-use rights to their generations, but a judgement that the generations are uncopyrightable or outright infringing on others copyright would make it effectively useless for the kinds of users who…

I wouldn't loose sleep over this if I was working for Midjourney. Copyright lobby is powerful, and when bad actors like Microsoft, Disney etc. jump onto the AI bandwagon and put their legal weight on their side of the lever, everything will turn out well (for them).

Re: Comparing Adobe Firefly, Dalle-2, and OpenJourney

#120
post #3

Midjurney is still so far ahead it's no competition. Did a lot of testing today and firefly generated so much errors with fingers and stuff, not seen that since the original stability release. Anyone know if the web firefly and the Photoshop version is the same model?

I share the same opinion, but also dislike these tests because each system benefits from a different approach to prompting. What I use to get a good result in MidJourney won't work in StableDiffusion for example. Instead when making these comparisons one needs to set an objective and have people who are familiar with each system to produce their nicest images - since this is a better reflection of the real world usage. For example, ask each participant to read a chapter/page from a book with a lot of specific imagery and then use AI to create what they think that looks like.

Regarding image generation in Photoshop I can confirm two things:

- It is excellent for in and out painting with a few exceptions*

- It remains poor for generating a brand new image

*Photoshop's generative fill is very good at extending landscapes, it will match lighting and according to the release video can be smart enough to observe what a reflection should contain even if that is not specifically included in the image (in their launch demo they showed how a reflection pool captured the underside of a vehicle.)

Where generative fill falls apart: Inserting new objects that are not well defined produces problems. Choosing something like a VW Beetle will produce a good result as it is well defined, choosing something like "boat", "dragon", or even "pirate's chest": will produce a range of images that do not necessarily fit the scene - this is likely because source imagery for such objects is likely vague and prone to different representations.

1st note about Firefly: Anything that is likely to produce a spherical looking shape tends to be blocked - likely because it resembles certain human anatomy. This is problematic when doing small touch ups such as fixing fingers.

A special note about photoshop versus other systems: Photoshop has the added problem of needing to match the resolution of the source material. Currently it achieves this from combining upscaling with resizing - this means that if one is extending an area with high detail, that detail cannot be maintained and instead is softer/blurrier than the original sections. It also means that if one extends directly from the border of an image, then a feathered edge becomes visible which must be corrected by hand.

I currently test the following AI generators, feel free to ask me about any of these: StableDiffusion (Automatic and InvokeAI), OpenAI's Dall-E 2, MidJourney, Stability AI's DreamStudio, and Adobe Firefly.

Post reply on HN