Live data from Hacker News

FLUX.1 Kontext

bfl.ai

81–90 of 140 posts

Re: FLUX.1 Kontext

#81

Some of these samples are rather cherry picked. Has anyone actually tried the professional headshot app of the "Kontext Apps"? https://replicate.com/flux-kontext-apps I've thrown half a dozen pictures of myself at it and it just completely replaced me with somebody else. To be fair, the final headshot does look very professional.

Nobody has solved the scientific problem of identity preservation for faces in one shot. Nobody has even solved hands.

I tried making a realistic image from a cartoon character but aged. It did very well, definitely recognisable as the same 'person'.

Re: FLUX.1 Kontext

#82

Some of these samples are rather cherry picked. Has anyone actually tried the professional headshot app of the "Kontext Apps"? https://replicate.com/flux-kontext-apps I've thrown half a dozen pictures of myself at it and it just completely replaced me with somebody else. To be fair, the final headshot does look very professional.

I tried a professional headshot prompt on the flux playground with a tired gym selfie and it kept it as myself, same expression, sweat, skin tone and all. It was like a background swap, then I expanded it to "make a professional headshot version of this image that would be good for social media, make the person smile, have a good pose and clothing, clean non-sweaty skin, etc" and it stayed pretty similar, except it swapped the clothing and gave me an awkward smile, which may be accurate for those kinds of things if you think about it.

Re: FLUX.1 Kontext

#84
post #57

Earlier quoted context omitted.

They chosen Asian traits that Western beauty standards fetishize that in Asia wouldn't be taken serious at all. I notice American text2image models tend to generate less attractive and more darker skinned humans where as Chinese text2image generate attractive and more light skinned humans. I think this is another area where Chinese AI models shine.

Wow, that is some straight-up overt racism. You should be ashamed.

It reads as racist if you parse it as (skin tone and attractiveness) but if you instead parse it as (skin tone) and (attractiveness), ie as two entirely unrelated characteristics of the output, then it reads as nothing more than a claim about relative differences in behavior between models.

Of course, given the sensitivity of the topic it is arguably somewhat inappropriate to make such observations without sufficient effort to clarify the precise meaning.

Re: FLUX.1 Kontext

#85

Earlier quoted context omitted.

Interesting, would you mind sharing? (imgur allows free image uploads, quick drag and drop) I do have a "works on my machine"* :) -- prompt "Model F keyboard", all settings disabled, on the smaller model, seems to have substantially more than no idea: https://imgur.com/a/32pV6Sp (Google Images comparison included to show in-the-wild "Model F keyboard", which may differ from my/your expected distribution) * my machine…

Your Google Images search indicates the original problem of models training on junk misinformation online. If AI scrapers are downloading every photo that's associated with "Model F Keyboard" like that, the models have no idea what is an IBM Model F, or its distinguishing characteristics, and what is some other company's, and what is misidentified. https://commons.wikimedia.org/wiki/Category:IBM_Model_F_Keyb... Speci…

> Eventually you may hit trademark protections. Reproducing a brand-name keyboard may be as difficult as simulating a celebrity's likeness.

Then the law is broken. Monetizing someone's likeness is an issue. Utilizing trademarked characteristics to promote your own product without permission is an issue. It's the downstream actions of the user that are the issue, not the ML model itself.

Models regurgitating copyrighted material verbatim is of course an entirely separate issue.

Re: FLUX.1 Kontext

#86

Earlier quoted context omitted.

Your Google Images search indicates the original problem of models training on junk misinformation online. If AI scrapers are downloading every photo that's associated with "Model F Keyboard" like that, the models have no idea what is an IBM Model F, or its distinguishing characteristics, and what is some other company's, and what is misidentified. https://commons.wikimedia.org/wiki/Category:IBM_Model_F_Keyb... Speci…

I agree with you vehemently. Another way of looking at it is, insistence on complete verisimilitude in an image generator is fundamentally in error. I would argue, even undesirable. I don't want to live in a world where a 45 year old keyboard that was only out for 4 years is readily imitated in every microscopic detail. I also find myself frustrated, and asking myself why. First thought that jumps in: it's very clear…

> if we're doing "the image generators don't get details right", it would seem to be there a lot simpler examples than OPs

There are different sorts of details though, and the distinctions are both useful and interesting to understanding the state of the art. If "man drinking coke" produces someone with 6 fingers holding a glass of water that's completely different from producing someone with 5 fingers holding a can of pepsi.

Notice that none of the images in your example got the function key placement correct. Clearly the model knows what a relatively modern keyboard is, and it even has some concept of a vaguely retro looking mechanical keyboard. However indeed I'm inclined to agree with OP that it has approximately zero idea what an "IBM model F" keyboard is. I'm not sure that's a failure of the model though - as you point out, it's an ancient and fairly obscure product.

Re: FLUX.1 Kontext

#87

Earlier quoted context omitted.

Your Google Images search indicates the original problem of models training on junk misinformation online. If AI scrapers are downloading every photo that's associated with "Model F Keyboard" like that, the models have no idea what is an IBM Model F, or its distinguishing characteristics, and what is some other company's, and what is misidentified. https://commons.wikimedia.org/wiki/Category:IBM_Model_F_Keyb... Speci…

> Eventually you may hit trademark protections. Reproducing a brand-name keyboard may be as difficult as simulating a celebrity's likeness. Then the law is broken. Monetizing someone's likeness is an issue. Utilizing trademarked characteristics to promote your own product without permission is an issue. It's the downstream actions of the user that are the issue, not the ML model itself. Models regurgitating copyright…

>> Eventually you may hit trademark protections.

> Then the law is broken.

> Utilizing trademarked characteristics to promote your own product without permission is an issue.

It sounds like you agree with parent that if your product reproduces trademark characteristics, it is utilizing trademarked characteristics. Just about at what layer you don't have responsibility. And the layer that has responsibility is the one that profits unjustly from the AI.

I'm interested if there's an argument for saying only the 2nd party user of the 1st party AI model, selling AI model output, to a 3rd party, is intuitively unfair.

I can't think of one. e.g. Disney launches some new cartoon or whatever. 1st party Openmetagoog, trains on it to make my "Video Episode Generator" product. Now, Openmetagoogs Community Pages are full of 30m video episodes made by their image generator. They didn't make them, nor do they promote them. Inuitively, Openmetagoog a competitor for manufacturing my IP, and that is also intuitively wrong. Your analysis would have us charge the users for sharing the output.

Re: FLUX.1 Kontext

#88

Earlier quoted context omitted.

> Eventually you may hit trademark protections. Reproducing a brand-name keyboard may be as difficult as simulating a celebrity's likeness. Then the law is broken. Monetizing someone's likeness is an issue. Utilizing trademarked characteristics to promote your own product without permission is an issue. It's the downstream actions of the user that are the issue, not the ML model itself. Models regurgitating copyright…

>> Eventually you may hit trademark protections. > Then the law is broken. > Utilizing trademarked characteristics to promote your own product without permission is an issue. It sounds like you agree with parent that if your product reproduces trademark characteristics, it is utilizing trademarked characteristics. Just about at what layer you don't have responsibility. And the layer that has responsibility is the one…

> if your product reproduces trademark characteristics, it is utilizing trademarked characteristics.

I wouldn't agree with that, no. To my mind "utilizing" generally requires intent at least in the context we're discussing here (ie moral or legal obligations). I'd remind you that the entire point of trademark is (approximately) to prevent brand confusion within the market.

> Your analysis would have us charge the users for sharing the output.

Precisely. I see it as both a matter of intent and concrete damages. Creating something (pencil, diffusion model, camera, etc) that could possibly be used in a manner that violates the law is not a problem. It is the end user violating the law that is at fault.

Imagine an online community that uses blender to create disney knockoffs and shares them publicly. Blender is not at fault and the creation of the knockoffs themselves (ie in private) is not the issue either. It's the part where the users proceed to publicly share them that poses the problem.

> They didn't make them, nor do they promote them.

By the same logic youtube neither creates nor promotes pirated content that gets uploaded. We have DMCA takedown notices for dealing with precisely this issue.

> Inuitively, Openmetagoog a competitor for manufacturing my IP, and that is also intuitively wrong.

Let's be clear about the distinction between trademark and copyright here. Outputting a verbatim copy is indeed a problem. Outputting a likeness is not, but an end user could certainly proceed to (mis)use that output in a manner that is.

Intent matters here. A product whose primary purpose is IP infringement is entirely different from one whose purpose is general but could potentially be used to infringe.

Re: FLUX.1 Kontext

#89
post #77
post #48

Earlier quoted context omitted.

It seems more accurate than 4o image generation in terms of preserving original details. If I give it my 3D animal character and ask it for a minor change like changing the lighting, 4o will completely mangle the face of my character, it will change the body and other details slightly. This Flux model keeps the visible geometry almost perfectly the same even when asked to significantly change the pose or lighting

gpt-image-1 (aka "4o") is still the most useful general purpose image model, but damn does this come close. I'm deep in this space and feel really good about FLUX.1 Kontext. It fills a much-needed gap, and it makes sure that OpenAI / Google aren't the runaway victors of images and video. Prior to gpt-image-1, the biggest problems in images were: - prompt adherence - generation quality - instructiveness (eg. "put the…

Your comment is def why we come to HN :)

Thanks for the detailed info

Post reply on HN