If this is still a diffusion model, I wonder how well does it compare with NanoBanana.
FLUX.2: Frontier Visual Intelligence
71–80 of 124 posts
Re: FLUX.2: Frontier Visual Intelligence
#72Earlier quoted context omitted.
I think it's more likely this is just a niche that Midjourney has occupied.
If Midjourney is a niche, then what is the broader market for AI image generation? Porn, obviously, though if you look at what's popular on civitai.com, a lot of it isn't photo-realistic. That might change as photo-realistic models are fully out of the uncanny valley. Presumably personalized advertising, but this isn't something we've seen much of yet. Maybe this is about to explode into the mainstream. Perhaps stock…
Re: FLUX.2: Frontier Visual Intelligence
#73Earlier quoted context omitted.
I heard a possibly unsubstantiated rumor that they had a major failed training run with the video model and canceled the project.
Makes no sense since they should have checkpoints earlier in the run that they could restart from and they should have regular checks that keep track if a model has exploded etc.
I'd take it with a grain of salt; these people are chainsaw jugglers and know what they're doing, so any sort of major hiccup was probably planned for. They'd have plan b and c, at a minimum, and be ready to switch - the work isn't deterministic, so you have to be ready for failures. (If you sense an imminent failure, don't grab the spinny part of the chainsaw, let it fall and move on.)
Re: FLUX.2: Frontier Visual Intelligence
#74Earlier quoted context omitted.
If Midjourney is a niche, then what is the broader market for AI image generation? Porn, obviously, though if you look at what's popular on civitai.com, a lot of it isn't photo-realistic. That might change as photo-realistic models are fully out of the uncanny valley. Presumably personalized advertising, but this isn't something we've seen much of yet. Maybe this is about to explode into the mainstream. Perhaps stock…
Well, as I said, if I type "cat", the most reasonable interpretation of that text string is a perfectly realistic cat. If I want an "illustration" I can type in "illustration of a cat". Though of course that's still quite unspecific. There are countless possible unrealistic styles for pictures (e.g. line art, manga, oil painting, vector art etc), and the reasonable thing is that the users should specify which of thes…
I think we'll probably need a few more hardware generations before it becomes feasible to use chatgpt 5 level models with integrated image generation. The underlying language model and its capabilities, the RL regime, and compute haven't caught up to the chat models yet, although nano-banana is certainly doing something right.
Re: FLUX.2: Frontier Visual Intelligence
#75Earlier quoted context omitted.
FYI: photorealism is art that imitates photos, and I see the term misused a lot both in comments and prompts (where you'll actually get subideal results if you say "photorealism" instead of describing the camera that "shot" it!)
I meant it here in the sense of "as indistinguishable from a photo as the model can make it".
I've heard chairs of animation departments say they feel like this puts film departments under them as a subset rather than the other way around. It's a funny twist of fate, given that the tables turned on them ages ago.
Photorealistic models are just learning the rules of camera optics and physics. In other "styles", the models learn how to draw Pixar shaded volumes, thick lines, or whatever rules and patterns and aesthetics you teach.
Different styles can reinforce one another across stylistic boundaries and mixed data sets can make the generalization better (at the cost of excelling in one domain).
"Real life", it seems, might just be a filter amongst many equally valid interpretations.
Re: FLUX.2: Frontier Visual Intelligence
#76Updating the GenAI comparison website is starting to feel a bit Sisyphean with all the new models coming out lately, but the results are in for the Flux 2 Pro Editing model! https://genai-showdown.specr.net/image-editing It scored slightly higher than BFL's Kontext model, coming in around the middle of the pack at 6 / 12 points. I’ll also be introducing an additional numerical metric soon, so we can add more nuance t…
If their new fancy model is only middle of the pack, and they're not as open source as the Chinese Qwen image models (or ByteDance / Alibaba / Lightricks video models), what's the point?
It's not just prompt adherence, the image quality of Flux models has been pretty bad. Plastic skin, inhumanely chiseled chins, that general faux "AI" aura.
Indeed, the Flux samples in your test suite that "pass" look God-awful. It might "pass" from a technical standpoint, but there's no way I'd choose Flux to solve my workflows. It looks bad.
(I wonder if they lack people on their data team with good aesthetic taste. It may be as simple as that.)
I think this company is struggling. They're pinned between Google and the Chinese. It's a tough, unenviable spot to be in.
I think a lot of the foundation model companies in media are having a really hard time: RunwayML, PikaLabs, LumaLabs. Some of them have pivoted hard away from solving media for everyone. I don't think they can beat the deep-pocketed hyperscalers or the Chinese ecosystem.
BFL just raised a massive round, so what do I know? I just can't help but feel that even though Runway raised similar money, they're struggling really hard now. And I would really not want to be fighting against Google who is already ahead in the game.
Re: FLUX.2: Frontier Visual Intelligence
#77Earlier quoted context omitted.
Makes no sense since they should have checkpoints earlier in the run that they could restart from and they should have regular checks that keep track if a model has exploded etc.
I didn't read "major failed training run" as in "the process crashed and we lost all data" but more like "After spending N weeks on training, we still didn't achieve our target(s)", which could be considered "failing" as well.
LTX's first model felt two years behind SOTA when it launched, but they viewed it as a success and kept going.
The investment initially is low and can scale with confidence.
BFL goes radio silent and then drops stuff. Now they're dropping stuff that is clearly middle of the pack.
Re: FLUX.2: Frontier Visual Intelligence
#78Great, especially that they still have an open-weight variant of this new model too. But what happened to their work on their unreleased SOTA video model? did it stop being SOTA, others got ahead, and they folded the project, or what? YT video about it: https://youtu.be/svIHNnM1Pa0?t=208 They even removed the page of that: https://bfl.ai/up-next/
As a startup, they pivoted and focused on image models (they are model providers, and image models often have more use cases than video models, not to mention they continue to have bigger image dataset moat, not video).
If they have so much data, then why do Flux model outputs look so God-awful bad?
They have plastic skin, weird chins, and have that "AI" aura. Not the good AI aura, mind you. The cheap automated YouTube video kind that you immediately skip.
Flux 2 seems to suffer from the exact same problems.
Midjourney is ancient. Their CEO is off trying to build a 3D volume and dating companion or some nonsense and leaving the product without guidance and much change. It almost feels abandoned. But even so, Midjourney has 10,000x better aesthetics despite having terrible prompt adherence and control. Midjourney images are dripping with magazine spread or Pulitzer aesthetics. It's why Zuckerberg went to them to license their model instead of quasi "open source" BFL.
Even SDXL looks better, and that's a literal dinosaur.
Most of the amazing things you see on social media either come from Midjourney or SDXL. To this day.
Re: FLUX.2: Frontier Visual Intelligence
#79I just finished my Flux 2 testing (focusing on the Pro variant here: https://replicate.com/black-forest-labs/flux-2-pro ). Overall, it's a tough sell to use Flux 2 over Nano Banana for the same use cases, but even if Nano Banana didn't exist it's only an iterative improvement over Flux 1.1 Pro. Some notes: - Running my nuanced Nano Banana prompts though Flux 2, Flux 2 definitely has better prompt adherence than Flux…
The fact that you have the possibility of running Flux locally might be enough of an argument to sway the balance for some cases. For example, if you've already set up a workflow and Google jacks up the price, or changes the API, you have no choice but to go along. If BFL does the same, you at least have the option of running locally.
I personally prefer Qwen's performance here. I'm waiting to see other folks' takes.
The Qwen folks are also a lot more transparent, spend time community building, and iterate on releases much more rapidly. In the open rather than behind closed doors.
I don't like how secretive BFL is.
Re: FLUX.2: Frontier Visual Intelligence
#80Earlier quoted context omitted.
If Midjourney is a niche, then what is the broader market for AI image generation? Porn, obviously, though if you look at what's popular on civitai.com, a lot of it isn't photo-realistic. That might change as photo-realistic models are fully out of the uncanny valley. Presumably personalized advertising, but this isn't something we've seen much of yet. Maybe this is about to explode into the mainstream. Perhaps stock…
> If Midjourney is a niche, then what is the broader market for AI image generation? Midjourney is one aesthetically pleasing data point in a wide spectrum of possibilities and market solutions. Creator economy is huge and is outgrowing Hollywood and the Music Industry combined. There's all sorts of use cases in marketing, corporate, internal comms. There are weird new markets. A lot of people simply subscribe to Mid…