Don’t understand the remove from face example. Without other pictures showing the persons face, it’s just using some stereotypical image, no?
The slideshow appears to be glitched on that first example. The input image has a snowflake covering most of her face.
FLUX.1 Kontext
111–120 of 140 posts
Re: FLUX.1 Kontext
#112Earlier quoted context omitted.
Seems implementation is straightforward (very similar to everyone else, HiDream-E1, ICEdit, DreamO etc.), the magic is on data curation (which details are lightly shared).
I haven't been following image generation models closely, at a high level is this new Flux model still diffusion based, or have they moved to block autoregressive (possibly with diffusion for upscaling) similar to 4o?
That's not the same as a diffusion model.
Here is a post about the difference that seems right at first glance: https://diffusionflow.github.io/
Re: FLUX.1 Kontext
#113Some of these samples are rather cherry picked. Has anyone actually tried the professional headshot app of the "Kontext Apps"? https://replicate.com/flux-kontext-apps I've thrown half a dozen pictures of myself at it and it just completely replaced me with somebody else. To be fair, the final headshot does look very professional.
Re: FLUX.1 Kontext
#114Some of these samples are rather cherry picked. Has anyone actually tried the professional headshot app of the "Kontext Apps"? https://replicate.com/flux-kontext-apps I've thrown half a dozen pictures of myself at it and it just completely replaced me with somebody else. To be fair, the final headshot does look very professional.
I tried a professional headshot prompt on the flux playground with a tired gym selfie and it kept it as myself, same expression, sweat, skin tone and all. It was like a background swap, then I expanded it to "make a professional headshot version of this image that would be good for social media, make the person smile, have a good pose and clothing, clean non-sweaty skin, etc" and it stayed pretty similar, except it s…
> Preserve Intentionally
> Specify what should stay the same: “while keeping the same facial features”
> Use “maintain the original composition” to preserve layout
> For background changes: “Change the background to a beach while keeping the person in the exact same position”
So while the marketing seems to paint a picture that it'll preserve things automatically, and kind of understand exactly what you want changed, it doesn't seem like that's the full truth. You need to instead be very specific about what you want to preserve.
Re: FLUX.1 Kontext
#115Re: FLUX.1 Kontext
#116Earlier quoted context omitted.
> I spent two days trying to train a LoRa customization on top of Flux 1 dev on Windows with my RTX 4090 but can’t make Windows is mostly the issue, to really take advantage, you will need linux.
Even using WSL2 with Ubuntu isn't good enough ?
The main thing is having 1. Good images with adequate captions and 2. Knowing what settings to use.
Number 2 is much harder because there’s a lot of bad information out there and the people who train a ton of Loras aren’t usually keen to share. Still, the various programs usually have some defaults that should be acceptable.
Re: FLUX.1 Kontext
#117First time native image gen was introduced in Gemini 1.5 Flash if I'm not wrong, and then OpenAI was released for 4o which took over the internet by Ghibli Art.
We have been getting good quality images from almost all image generators like Midjourney, OpenAI and other providers, but the thing that made it special was true "multimodal" nature of it. Here's what I mean
When you used to ask chatgpt to create an image, it will rephrase that prompt and internally send that prompt to Dalle, similarly gemini would send it to Imagen which were diffusion models and they had little to know context in your next response about what's there in the previous image
In native image generation, it understands Audio, Text and even Image tokens in the same model and need not to rely on diffusion models internally, I don't think both Openai and google has released how they've trained it but my guess is that it's partially auto-regressive and diffusion but not sure about it
Re: FLUX.1 Kontext
#118Earlier quoted context omitted.
Founder of Replicate here. We should be on par or faster for all the top models. e.g. we have the fastest FLUX[dev]: https://artificialanalysis.ai/text-to-image/model-family/flu... If something's not as fast let me know and we can fix it. ben@replicate.com
Hey Ben, thanks for participating in this thread. And certainly also for all you and your team have built. Totally frank and possibly awkward question, you don't have to answer: how do you feel about a16z investing in everyone in this space? They invested in you. They're investing in your direct competitors (Fal, et al.) They're picking your downmarket and upmarket (Krea, et al.) They're picking consumer (Viggle, et…
Re: FLUX.1 Kontext
#119Earlier quoted context omitted.
It seems more accurate than 4o image generation in terms of preserving original details. If I give it my 3D animal character and ask it for a minor change like changing the lighting, 4o will completely mangle the face of my character, it will change the body and other details slightly. This Flux model keeps the visible geometry almost perfectly the same even when asked to significantly change the pose or lighting
gpt-image-1 (aka "4o") is still the most useful general purpose image model, but damn does this come close. I'm deep in this space and feel really good about FLUX.1 Kontext. It fills a much-needed gap, and it makes sure that OpenAI / Google aren't the runaway victors of images and video. Prior to gpt-image-1, the biggest problems in images were: - prompt adherence - generation quality - instructiveness (eg. "put the…
Glad to see this release. It also puts more pressure onto OpenAI to make their model less lobotomized and to increase its output quality. This is good for everyone.
Re: FLUX.1 Kontext
#120Earlier quoted context omitted.
They chosen Asian traits that Western beauty standards fetishize that in Asia wouldn't be taken serious at all. I notice American text2image models tend to generate less attractive and more darker skinned humans where as Chinese text2image generate attractive and more light skinned humans. I think this is another area where Chinese AI models shine.
Wow, that is some straight-up overt racism. You should be ashamed.
This particular woman looks Vietnamese to me, but I agree nothing about her appearance looks like anyone's fashion I know. But I only know California ABGs so that doesn't mean much.