Updating the GenAI comparison website is starting to feel a bit Sisyphean with all the new models coming out lately, but the results are in for the Flux 2 Pro Editing model! https://genai-showdown.specr.net/image-editing It scored slightly higher than BFL's Kontext model, coming in around the middle of the pack at 6 / 12 points. I’ll also be introducing an additional numerical metric soon, so we can add more nuance t…
Clearly Google is winning this by some margin Seedream is also very good and makes me think the next version will challenge Google for SOTA image gen Increasingly feels like image gen is a solved problem
FLUX.2: Frontier Visual Intelligence
101–110 of 124 posts
Re: FLUX.2: Frontier Visual Intelligence
#102Updating the GenAI comparison website is starting to feel a bit Sisyphean with all the new models coming out lately, but the results are in for the Flux 2 Pro Editing model! https://genai-showdown.specr.net/image-editing It scored slightly higher than BFL's Kontext model, coming in around the middle of the pack at 6 / 12 points. I’ll also be introducing an additional numerical metric soon, so we can add more nuance t…
Clearly Google is winning this by some margin Seedream is also very good and makes me think the next version will challenge Google for SOTA image gen Increasingly feels like image gen is a solved problem
Also it doesn't feel solved to me at all. There is no general model, perhaps it cannot reasonably exist. I think these tests are benchmarks are smart, but they don't show the whole picture.
Domain specific image generation tasks still require a domain specific models. For art purposes SD1.5 with specialized and finely tuned checkpoints will still provide the best results by far. It is also limited, but I think it dampened the hype for new image generators significantly.
Re: FLUX.2: Frontier Visual Intelligence
#103Updating the GenAI comparison website is starting to feel a bit Sisyphean with all the new models coming out lately, but the results are in for the Flux 2 Pro Editing model! https://genai-showdown.specr.net/image-editing It scored slightly higher than BFL's Kontext model, coming in around the middle of the pack at 6 / 12 points. I’ll also be introducing an additional numerical metric soon, so we can add more nuance t…
Re: FLUX.2: Frontier Visual Intelligence
#104Updating the GenAI comparison website is starting to feel a bit Sisyphean with all the new models coming out lately, but the results are in for the Flux 2 Pro Editing model! https://genai-showdown.specr.net/image-editing It scored slightly higher than BFL's Kontext model, coming in around the middle of the pack at 6 / 12 points. I’ll also be introducing an additional numerical metric soon, so we can add more nuance t…
The comparison are very useful but also quite limited in terms of styles. Models tend to have extremely diverse abilities in following a given style against steering to its own. It's pretty obvious that OpenAI is terrible at it -- it is known for its unmissable touch. However, for Flux it really depends on the style. They already posted at some point that they changed their training to avoid averaging different style…
Generative: https://genai-showdown.specr.net
Editing: https://genai-showdown.specr.net/image-editing
Style is mostly irrelevant for editing, since the goal is to integrate seamlessly with the existing image. The focus is on performing relatively surgical edits or modifications to existing imagery while minimizing changes to the rest of the image. It is also primarily concerned with realism, though there are some illustrative examples (the JAWS poster, Great Wave off Kanagawa).
This contrasts with the generative section though even then the emphasis is on prompt adherence, and style/fidelity take a backseat (which is honestly what 99% of existing generative benchmarks already focus on).
Re: FLUX.2: Frontier Visual Intelligence
#105Earlier quoted context omitted.
How much energy does BFL have to keep playing this game against Google and ByteDance (SeeDream)? If their new fancy model is only middle of the pack, and they're not as open source as the Chinese Qwen image models (or ByteDance / Alibaba / Lightricks video models), what's the point? It's not just prompt adherence, the image quality of Flux models has been pretty bad. Plastic skin, inhumanely chiseled chins, that gene…
i may be wrong, but it doesn't seem like BFL is struggling to me. they were apparently founded in august 2024, and have already signed $100M+ revenue deals with customers like meta ( https://www.bloomberg.com/news/articles/2025-09-09/meta-to-p... ) in fact, it seems like BFL has benefited a lot by becoming the go-to alternative for big enterprise customers who don't want to be dependent on google
Re: FLUX.2: Frontier Visual Intelligence
#106Earlier quoted context omitted.
Wow, I didn't hear about this. That's impressive, and kudos to the team. That's why they raised the massive round, then. But this just leads to more questions - I have to wonder if and for how long this is just going to be to plug in a gap for Meta's own AI product offering. At some point they'll want to build their own in-house models or perhaps just acquire BFL. Zuckerberg would not be printing AI data centers if t…
idk it seems pretty clear BFL’s target market is developers not graphic designers. and for developers at scale like Meta and Adobe, it’s pretty incredible a tiny startup like BFL has become the primary alternative to Google with 1/100th of the resources within 12 months of their founding, doing hundreds of millions of revenue the Chinese models are great, but no serious enterprise developer is going to bet their imag…
Re: FLUX.2: Frontier Visual Intelligence
#107> Launch Partners Wow, the Krea relationship soured? These are both a16z companies and they've worked on private model development before. Krea.1 was supposed to be something to compete with Midjourney aesthetics and get away from the plastic-y Flux models with artificial skin tones, weird chins, etc. This list of partners includes all of Krea's competitors: HiggsField (current aggregator leader), Freepik, "Open"Art,…
They messed up. We (Krea) were also surprised. They put our logo after we pointed it out. Nice eye!
Re: FLUX.2: Frontier Visual Intelligence
#108Earlier quoted context omitted.
Clearly Google is winning this by some margin Seedream is also very good and makes me think the next version will challenge Google for SOTA image gen Increasingly feels like image gen is a solved problem
I think the margin isn't that large to be honest. If we compare available resources and data it is quite tiny and perhaps should be larger. Also it doesn't feel solved to me at all. There is no general model, perhaps it cannot reasonably exist. I think these tests are benchmarks are smart, but they don't show the whole picture. Domain specific image generation tasks still require a domain specific models. For art pur…
I understand most outputs could be fine tuned for most domains, but still felt sd1.5 had a resolution ceiling, and a complexity ceiling no matter how good the fine tuning
Re: FLUX.2: Frontier Visual Intelligence
#109Updating the GenAI comparison website is starting to feel a bit Sisyphean with all the new models coming out lately, but the results are in for the Flux 2 Pro Editing model! https://genai-showdown.specr.net/image-editing It scored slightly higher than BFL's Kontext model, coming in around the middle of the pack at 6 / 12 points. I’ll also be introducing an additional numerical metric soon, so we can add more nuance t…
Re: FLUX.2: Frontier Visual Intelligence
#110Earlier quoted context omitted.
The comparison are very useful but also quite limited in terms of styles. Models tend to have extremely diverse abilities in following a given style against steering to its own. It's pretty obvious that OpenAI is terrible at it -- it is known for its unmissable touch. However, for Flux it really depends on the style. They already posted at some point that they changed their training to avoid averaging different style…
The site is broken up into "Editing Comparison" and a "Generative Comparison" sections. Generative: https://genai-showdown.specr.net Editing: https://genai-showdown.specr.net/image-editing Style is mostly irrelevant for editing, since the goal is to integrate seamlessly with the existing image. The focus is on performing relatively surgical edits or modifications to existing imagery while minimizing changes to the re…
If you look for example at "Mermaid Disciplinary Committee", every single image is in a very different style, each that you can consider a default of what the model assume would be for the specific prompt. It's quite obvious that these styles were 'baked in' the models, and it's not clear how much you can steer in a specific style. If you look at "The Yarrctic Circle", a lot more models default to a kind of "generic concept art" style (the "by greg rutkowski" meme) but even then I would classify the results as at least 5 distinct styles. So for me this benchmark is not checking style at all, unless you consider style to be just around 4 categories (cartoon, anime, realistic, painterly).
So regarding image editing, I did my own tests at the first release of Flux tools, and found that it was almost impossible to get any decent results on some specific styles, specifically cartoon and concept art styles. I think the tools focus on what imaginary marketing people would want (like "put this can of sugary beverage into an idyllic scene") rather than such use cases. So editing like "color this" or other changes would just be terrible, and certainly unusable.