Live data from Hacker News

FLUX.1 Kontext

bfl.ai

111–120 of 140 posts

Re: FLUX.1 Kontext

#111

Don’t understand the remove from face example. Without other pictures showing the persons face, it’s just using some stereotypical image, no?

The slideshow appears to be glitched on that first example. The input image has a snowflake covering most of her face.

That's the point, it can remove it.

Re: FLUX.1 Kontext

#112
post #22

Earlier quoted context omitted.

Seems implementation is straightforward (very similar to everyone else, HiDream-E1, ICEdit, DreamO etc.), the magic is on data curation (which details are lightly shared).

I haven't been following image generation models closely, at a high level is this new Flux model still diffusion based, or have they moved to block autoregressive (possibly with diffusion for upscaling) similar to 4o?

Well it's a "generative flow matching model"

That's not the same as a diffusion model.

Here is a post about the difference that seems right at first glance: https://diffusionflow.github.io/

Re: FLUX.1 Kontext

#113

Some of these samples are rather cherry picked. Has anyone actually tried the professional headshot app of the "Kontext Apps"? https://replicate.com/flux-kontext-apps I've thrown half a dozen pictures of myself at it and it just completely replaced me with somebody else. To be fair, the final headshot does look very professional.

It's convenient but the results are def not significantly better than available free stuff

Re: FLUX.1 Kontext

#114
post #82

Some of these samples are rather cherry picked. Has anyone actually tried the professional headshot app of the "Kontext Apps"? https://replicate.com/flux-kontext-apps I've thrown half a dozen pictures of myself at it and it just completely replaced me with somebody else. To be fair, the final headshot does look very professional.

I tried a professional headshot prompt on the flux playground with a tired gym selfie and it kept it as myself, same expression, sweat, skin tone and all. It was like a background swap, then I expanded it to "make a professional headshot version of this image that would be good for social media, make the person smile, have a good pose and clothing, clean non-sweaty skin, etc" and it stayed pretty similar, except it s…

It isn't mentioned on https://replicate.com/flux-kontext-apps/professional-headsho..., but on https://replicate.com/black-forest-labs/flux-kontext-pro under the "Prompting Best Practices" section is says this:

> Preserve Intentionally

> Specify what should stay the same: “while keeping the same facial features”

> Use “maintain the original composition” to preserve layout

> For background changes: “Change the background to a beach while keeping the person in the exact same position”

So while the marketing seems to paint a picture that it'll preserve things automatically, and kind of understand exactly what you want changed, it doesn't seem like that's the full truth. You need to instead be very specific about what you want to preserve.

Re: FLUX.1 Kontext

#115
post #70
post #39

Earlier quoted context omitted.

Well thank you I will test that

SimpleTuner is dependant on Microsoft's DeepSpeed which doesnt work on Windows :) So you probably better off using Ai-ToolKit https://github.com/ostris/ai-toolkit

OneTrainer would be another “easy” option.

Re: FLUX.1 Kontext

#116
post #37
post #32

Earlier quoted context omitted.

> I spent two days trying to train a LoRa customization on top of Flux 1 dev on Windows with my RTX 4090 but can’t make Windows is mostly the issue, to really take advantage, you will need linux.

Even using WSL2 with Ubuntu isn't good enough ?

Nah, that’s fine. So is Windows for most tools.

The main thing is having 1. Good images with adequate captions and 2. Knowing what settings to use.

Number 2 is much harder because there’s a lot of bad information out there and the people who train a ton of Loras aren’t usually keen to share. Still, the various programs usually have some defaults that should be acceptable.

Re: FLUX.1 Kontext

#117
So here is my understanding of current native image generation scenario, I might be wrong so please correct me, I'm still learning it and I'd appreaciate the help.

First time native image gen was introduced in Gemini 1.5 Flash if I'm not wrong, and then OpenAI was released for 4o which took over the internet by Ghibli Art.

We have been getting good quality images from almost all image generators like Midjourney, OpenAI and other providers, but the thing that made it special was true "multimodal" nature of it. Here's what I mean

When you used to ask chatgpt to create an image, it will rephrase that prompt and internally send that prompt to Dalle, similarly gemini would send it to Imagen which were diffusion models and they had little to know context in your next response about what's there in the previous image

In native image generation, it understands Audio, Text and even Image tokens in the same model and need not to rely on diffusion models internally, I don't think both Openai and google has released how they've trained it but my guess is that it's partially auto-regressive and diffusion but not sure about it

Re: FLUX.1 Kontext

#118
post #79
post #78

Earlier quoted context omitted.

Founder of Replicate here. We should be on par or faster for all the top models. e.g. we have the fastest FLUX[dev]: https://artificialanalysis.ai/text-to-image/model-family/flu... If something's not as fast let me know and we can fix it. ben@replicate.com

Hey Ben, thanks for participating in this thread. And certainly also for all you and your team have built. Totally frank and possibly awkward question, you don't have to answer: how do you feel about a16z investing in everyone in this space? They invested in you. They're investing in your direct competitors (Fal, et al.) They're picking your downmarket and upmarket (Krea, et al.) They're picking consumer (Viggle, et…

That feels like the VC equivalent of buying a market-specific fund, so fairly par for the course?

Re: FLUX.1 Kontext

#119
post #77
post #48

Earlier quoted context omitted.

It seems more accurate than 4o image generation in terms of preserving original details. If I give it my 3D animal character and ask it for a minor change like changing the lighting, 4o will completely mangle the face of my character, it will change the body and other details slightly. This Flux model keeps the visible geometry almost perfectly the same even when asked to significantly change the pose or lighting

gpt-image-1 (aka "4o") is still the most useful general purpose image model, but damn does this come close. I'm deep in this space and feel really good about FLUX.1 Kontext. It fills a much-needed gap, and it makes sure that OpenAI / Google aren't the runaway victors of images and video. Prior to gpt-image-1, the biggest problems in images were: - prompt adherence - generation quality - instructiveness (eg. "put the…

When I first saw gpt-image-1 I was equally scared that OpenAI had used its resources to push so far ahead that more open models would be left completely in the dust for the significant future.

Glad to see this release. It also puts more pressure onto OpenAI to make their model less lobotomized and to increase its output quality. This is good for everyone.

Re: FLUX.1 Kontext

#120
post #57

Earlier quoted context omitted.

They chosen Asian traits that Western beauty standards fetishize that in Asia wouldn't be taken serious at all. I notice American text2image models tend to generate less attractive and more darker skinned humans where as Chinese text2image generate attractive and more light skinned humans. I think this is another area where Chinese AI models shine.

Wow, that is some straight-up overt racism. You should be ashamed.

Asians can be pretty colorist within themselves and they're not going to listen to you when you tell them it's bad. Asian women love skin-lightening creams.

This particular woman looks Vietnamese to me, but I agree nothing about her appearance looks like anyone's fashion I know. But I only know California ABGs so that doesn't mean much.

Post reply on HN