Live data from Hacker News

Flux 3

bfl.ai

31–40 of 151 posts

Re: Flux 3

#31
post #28

I hope the open-weight versions will be SOTA. > Over the next few weeks and months, we will make the following capabilities available > Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”) > We will also release more technical details on the underlying approach.

Open-weight promise seems nice, but 1) I've seen people commenting that these hopes amounted to nothing for some previous releases (no idea which promises were made though); and 2) if the backbone is released as Dev, what will be missing? I can't easily tell from the post.

dev variants are usually cfg distilled which means that directly finetuning isn’t as effective. In the past, for the flux2 klein models,they released base versions that are not distilled. So it will probably be a while before you can fully take advantage of the open weights.

Re: Flux 3

#32
i don't get why they are investing money on image/video gen. All generations i have see looked blurry, lacking in fine details and missing the artistic touch (lacks meaning? lifeless?)

Re: Flux 3

#33

It's incredible how negative and dismissive the comments in here are while here I am thinking the model actually looks impressively capable. But then again I heard the downers have always been the first to leave their dung comments here so let's see...

It's possible it's because some people care about the consequences of what we're doing as a society beyond the mindless satiation of curiosity. Nothing wrong with curiosity in my books by the way, but I do think isolating it the way we've done in the technical fields is a dangerous and irresponsible attitude.

Re: Flux 3

#34
I have a feeling open-weight models ought to be outperforming proprietary ones by now, but that still hasn’t happened. So far, Nano Banana and GPT-2 Image seem to be the best in class, and Flux still isn’t crossing that quality bar.

Re: Flux 3

#36

- Showed close to zero examples of people. - Frivolous use of the term World Model. - Claims 20 seconds of video, shows only jumpcuts. Coming soon!

Yeah, very weird launch. I was like... where's the videos?

Re: Flux 3

#37
post #3

Lots of words about multi-modal but then this: > our mission to develop real-world visual intelligence Visual is mono-modal, isn't it?

its doing video, audio, images and motion. I think that counts as multimodal.

Re: Flux 3

#38

- Showed close to zero examples of people. - Frivolous use of the term World Model. - Claims 20 seconds of video, shows only jumpcuts. Coming soon!

> Frivolous use of the term World Model

The term "world model" as it was once used in model-based RL can now apparently refer to anything as silly as linear regression. Then again, the RL folks probably borrowed the term from behavioral scientists before them. It's probably best to simply accept this :/

A similar thing happened to "object oriented" which has been misused by philosophers and visual artists alike.

Re: Flux 3

#39
post #32

i don't get why they are investing money on image/video gen. All generations i have see looked blurry, lacking in fine details and missing the artistic touch (lacks meaning? lifeless?)

> why they are investing money

Maybe they presume that after a series of "good enough to some" they may be getting near the Real Thing?

Post reply on HN