Some of these samples are rather cherry picked. Has anyone actually tried the professional headshot app of the "Kontext Apps"? https://replicate.com/flux-kontext-apps I've thrown half a dozen pictures of myself at it and it just completely replaced me with somebody else. To be fair, the final headshot does look very professional.
Nobody has solved the scientific problem of identity preservation for faces in one shot. Nobody has even solved hands.
FLUX.1 Kontext
81–90 of 140 posts
Re: FLUX.1 Kontext
#82Some of these samples are rather cherry picked. Has anyone actually tried the professional headshot app of the "Kontext Apps"? https://replicate.com/flux-kontext-apps I've thrown half a dozen pictures of myself at it and it just completely replaced me with somebody else. To be fair, the final headshot does look very professional.
Re: FLUX.1 Kontext
#83Re: FLUX.1 Kontext
#84Earlier quoted context omitted.
They chosen Asian traits that Western beauty standards fetishize that in Asia wouldn't be taken serious at all. I notice American text2image models tend to generate less attractive and more darker skinned humans where as Chinese text2image generate attractive and more light skinned humans. I think this is another area where Chinese AI models shine.
Wow, that is some straight-up overt racism. You should be ashamed.
Of course, given the sensitivity of the topic it is arguably somewhat inappropriate to make such observations without sufficient effort to clarify the precise meaning.
Re: FLUX.1 Kontext
#85Earlier quoted context omitted.
Interesting, would you mind sharing? (imgur allows free image uploads, quick drag and drop) I do have a "works on my machine"* :) -- prompt "Model F keyboard", all settings disabled, on the smaller model, seems to have substantially more than no idea: https://imgur.com/a/32pV6Sp (Google Images comparison included to show in-the-wild "Model F keyboard", which may differ from my/your expected distribution) * my machine…
Your Google Images search indicates the original problem of models training on junk misinformation online. If AI scrapers are downloading every photo that's associated with "Model F Keyboard" like that, the models have no idea what is an IBM Model F, or its distinguishing characteristics, and what is some other company's, and what is misidentified. https://commons.wikimedia.org/wiki/Category:IBM_Model_F_Keyb... Speci…
Then the law is broken. Monetizing someone's likeness is an issue. Utilizing trademarked characteristics to promote your own product without permission is an issue. It's the downstream actions of the user that are the issue, not the ML model itself.
Models regurgitating copyrighted material verbatim is of course an entirely separate issue.
Re: FLUX.1 Kontext
#86Earlier quoted context omitted.
Your Google Images search indicates the original problem of models training on junk misinformation online. If AI scrapers are downloading every photo that's associated with "Model F Keyboard" like that, the models have no idea what is an IBM Model F, or its distinguishing characteristics, and what is some other company's, and what is misidentified. https://commons.wikimedia.org/wiki/Category:IBM_Model_F_Keyb... Speci…
I agree with you vehemently. Another way of looking at it is, insistence on complete verisimilitude in an image generator is fundamentally in error. I would argue, even undesirable. I don't want to live in a world where a 45 year old keyboard that was only out for 4 years is readily imitated in every microscopic detail. I also find myself frustrated, and asking myself why. First thought that jumps in: it's very clear…
There are different sorts of details though, and the distinctions are both useful and interesting to understanding the state of the art. If "man drinking coke" produces someone with 6 fingers holding a glass of water that's completely different from producing someone with 5 fingers holding a can of pepsi.
Notice that none of the images in your example got the function key placement correct. Clearly the model knows what a relatively modern keyboard is, and it even has some concept of a vaguely retro looking mechanical keyboard. However indeed I'm inclined to agree with OP that it has approximately zero idea what an "IBM model F" keyboard is. I'm not sure that's a failure of the model though - as you point out, it's an ancient and fairly obscure product.
Re: FLUX.1 Kontext
#87Earlier quoted context omitted.
Your Google Images search indicates the original problem of models training on junk misinformation online. If AI scrapers are downloading every photo that's associated with "Model F Keyboard" like that, the models have no idea what is an IBM Model F, or its distinguishing characteristics, and what is some other company's, and what is misidentified. https://commons.wikimedia.org/wiki/Category:IBM_Model_F_Keyb... Speci…
> Eventually you may hit trademark protections. Reproducing a brand-name keyboard may be as difficult as simulating a celebrity's likeness. Then the law is broken. Monetizing someone's likeness is an issue. Utilizing trademarked characteristics to promote your own product without permission is an issue. It's the downstream actions of the user that are the issue, not the ML model itself. Models regurgitating copyright…
> Then the law is broken.
> Utilizing trademarked characteristics to promote your own product without permission is an issue.
It sounds like you agree with parent that if your product reproduces trademark characteristics, it is utilizing trademarked characteristics. Just about at what layer you don't have responsibility. And the layer that has responsibility is the one that profits unjustly from the AI.
I'm interested if there's an argument for saying only the 2nd party user of the 1st party AI model, selling AI model output, to a 3rd party, is intuitively unfair.
I can't think of one. e.g. Disney launches some new cartoon or whatever. 1st party Openmetagoog, trains on it to make my "Video Episode Generator" product. Now, Openmetagoogs Community Pages are full of 30m video episodes made by their image generator. They didn't make them, nor do they promote them. Inuitively, Openmetagoog a competitor for manufacturing my IP, and that is also intuitively wrong. Your analysis would have us charge the users for sharing the output.
Re: FLUX.1 Kontext
#88Earlier quoted context omitted.
> Eventually you may hit trademark protections. Reproducing a brand-name keyboard may be as difficult as simulating a celebrity's likeness. Then the law is broken. Monetizing someone's likeness is an issue. Utilizing trademarked characteristics to promote your own product without permission is an issue. It's the downstream actions of the user that are the issue, not the ML model itself. Models regurgitating copyright…
>> Eventually you may hit trademark protections. > Then the law is broken. > Utilizing trademarked characteristics to promote your own product without permission is an issue. It sounds like you agree with parent that if your product reproduces trademark characteristics, it is utilizing trademarked characteristics. Just about at what layer you don't have responsibility. And the layer that has responsibility is the one…
I wouldn't agree with that, no. To my mind "utilizing" generally requires intent at least in the context we're discussing here (ie moral or legal obligations). I'd remind you that the entire point of trademark is (approximately) to prevent brand confusion within the market.
> Your analysis would have us charge the users for sharing the output.
Precisely. I see it as both a matter of intent and concrete damages. Creating something (pencil, diffusion model, camera, etc) that could possibly be used in a manner that violates the law is not a problem. It is the end user violating the law that is at fault.
Imagine an online community that uses blender to create disney knockoffs and shares them publicly. Blender is not at fault and the creation of the knockoffs themselves (ie in private) is not the issue either. It's the part where the users proceed to publicly share them that poses the problem.
> They didn't make them, nor do they promote them.
By the same logic youtube neither creates nor promotes pirated content that gets uploaded. We have DMCA takedown notices for dealing with precisely this issue.
> Inuitively, Openmetagoog a competitor for manufacturing my IP, and that is also intuitively wrong.
Let's be clear about the distinction between trademark and copyright here. Outputting a verbatim copy is indeed a problem. Outputting a likeness is not, but an end user could certainly proceed to (mis)use that output in a manner that is.
Intent matters here. A product whose primary purpose is IP infringement is entirely different from one whose purpose is general but could potentially be used to infringe.
Re: FLUX.1 Kontext
#89Earlier quoted context omitted.
It seems more accurate than 4o image generation in terms of preserving original details. If I give it my 3D animal character and ask it for a minor change like changing the lighting, 4o will completely mangle the face of my character, it will change the body and other details slightly. This Flux model keeps the visible geometry almost perfectly the same even when asked to significantly change the pose or lighting
gpt-image-1 (aka "4o") is still the most useful general purpose image model, but damn does this come close. I'm deep in this space and feel really good about FLUX.1 Kontext. It fills a much-needed gap, and it makes sure that OpenAI / Google aren't the runaway victors of images and video. Prior to gpt-image-1, the biggest problems in images were: - prompt adherence - generation quality - instructiveness (eg. "put the…
Thanks for the detailed info