Live data from Hacker News

Generative AI Image Editing Showdown

genai-showdown.specr.net

71–80 of 83 posts

Re: Generative AI Image Editing Showdown

#71

Everyone is sleeping on Gemini 2.5 Flash Image / Nano Banana. As shown in the OP, it's substantially more powerful than most other models while at the same price-per-image, and due to its text encoder it can handle significantly larger and more nuanced prompts to get exactly what you want. I open-sourced a Python package for generating from it with examples ( https://github.com/minimaxir/gemimg ) and am currently wor…

half the time when i try to use nano banana, AI Studio fails, telling me it can't generate for some unspecified reason. these aren't cases where I'm trying to do something that skirts the edge of copyright, either (like "Ghiblifying" images, for example). that said, when it does work, it is super impressive.

Let's just say I've tested around this.

Copyright: Zero guardrails on anything related to third-party IP, which lets you do some funny things. (I'm including a picture/prompt of Super Mario, Mickey Mouse, and Bugs Bunny partying at a nightclub in the blog post)

Moderation: It has far fewer guardrails and any other Google AI product I've tried, and it is possible to prompt engineer some images that would definitely be considered NSFW by most people — more NSFW than actual NSFW image generators (a post-generation filter will catch most nudity, however). I have not had any rejections for more innocous queries that could be misinterpreted as being NSFW.

Re: Generative AI Image Editing Showdown

#72

Earlier quoted context omitted.

half the time when i try to use nano banana, AI Studio fails, telling me it can't generate for some unspecified reason. these aren't cases where I'm trying to do something that skirts the edge of copyright, either (like "Ghiblifying" images, for example). that said, when it does work, it is super impressive.

It might be the safety moderation system. It's rather aggressive and when it does kick in (at least in the API), it often returns an empty response giving basically zero indication as to the root cause.

The empty response issue is annoying since there is already a PROHIBITED_CONTENT flag, but it is not used in this case.

Re: Generative AI Image Editing Showdown

#73

I still feel varying the prompt text, number of tries, and varying strictness combined with only showing the result most liked dilute most of the value in these test. It would be better if there was one prompt 8/10 human editors understood and implemented correctly and then every model got 5 generation attempts with that exact prompt on different seeds or something. If it were about "who can create the best image wit…

The OpenAI gpt-image-1 example was supposed to be noted as for the "You Only Move Twice" test.

Re: Generative AI Image Editing Showdown

#74
post #41

Earlier quoted context omitted.

You have to REALLY be into AI to do this for generation/API cost reasons (or willing to have this as a hacking project of the month expense). Even ignoring electricity, a 16 GB 5060 Ti is more expensive than 16,000 image generations. Assuming you do one every 15 seconds, that's 240,000 seconds -> more than 2 months of usage at an hour a day of generations. If you've already got a decent GPU (or were going to get one…

GPUs are needed for plenty of reasons. I assume plenty have a decent dGPU, even on laptops.

The point is just having a "decent" dGPU isn't enough. Even at 16 GB you're already quantizing Flux pretty heavily, someone with a 4080 gaming laptop is going to be disappointed trying to work with 12 GB.

Re: Generative AI Image Editing Showdown

#75

Earlier quoted context omitted.

I understood that much, at least from the description you added on the Kontext result. I agree that you should provide more information here, though, especially around "we adapt the exact prompts depending on the model", since your strategy here could also reflect model strengths and weaknesses.

Good point! Perhaps I should add in the "final model-specific prompt", or place them in an errata section.

[deleted]

Re: Generative AI Image Editing Showdown

#76

Earlier quoted context omitted.

I understood that much, at least from the description you added on the Kontext result. I agree that you should provide more information here, though, especially around "we adapt the exact prompts depending on the model", since your strategy here could also reflect model strengths and weaknesses.

Good point! Perhaps I should add in the "final model-specific prompt", or place them in an errata section.

By the way, this is what I got from Kontext after just a couple of tries: https://i.imgur.com/J4LwkVI.png

Prompt: "Keeping the glass and the hand behind the glass the same, please change only the three brown candies in the glass into green, yellow, red, and orange candies. Make no other changes. Change the reflection to remove the brown candy too." Seed was 1070229954903864, but your setup is probably too different for that to help.

It seems like Gemini 2.5 Flash was the only model that successfully removed the reflections...it should get some points for that!

Re: Generative AI Image Editing Showdown

#77

Kontext is very good. Get yourself a 5060 ti 16GB and never have to pay for API calls again for this purpose, at least not when you have the time spare. If you need this sort of editing at the speed of gui-clicking + 10s, then you'll need to pay API tolls, or capex for > 5070/80.

You have to REALLY be into AI to do this for generation/API cost reasons (or willing to have this as a hacking project of the month expense). Even ignoring electricity, a 16 GB 5060 Ti is more expensive than 16,000 image generations. Assuming you do one every 15 seconds, that's 240,000 seconds -> more than 2 months of usage at an hour a day of generations. If you've already got a decent GPU (or were going to get one…

>a 16 GB 5060 Ti is more expensive than 16,000 image generations

Sure, but now you get a good gaming GPU that you can write off as a business expense.

Re: Generative AI Image Editing Showdown

#78

Kontext is very good. Get yourself a 5060 ti 16GB and never have to pay for API calls again for this purpose, at least not when you have the time spare. If you need this sort of editing at the speed of gui-clicking + 10s, then you'll need to pay API tolls, or capex for > 5070/80.

You have to REALLY be into AI to do this for generation/API cost reasons (or willing to have this as a hacking project of the month expense). Even ignoring electricity, a 16 GB 5060 Ti is more expensive than 16,000 image generations. Assuming you do one every 15 seconds, that's 240,000 seconds -> more than 2 months of usage at an hour a day of generations. If you've already got a decent GPU (or were going to get one…

16,000? Where are buying your GPU, or API calls? If you don’t want to wait for a bargain then $450 will get you the GPU, and even at that price you’d only be able to buy about 10,000 standard-resolution image gen api calls. Do you do design? Editing? Touch up? You can easily blow through a few hundred api calls an hour: “Turn the stitching green… slightly less saturated… now make the stitches more ragged… a little more… now just slightly less”.

Clearly you’re looking at the task through the eyes of a hobbyist or “of the month” project so the workflow and pace may not be obvious but API budgets spend fast. Just look at the benchmarks in this article to see how many tried some of these changes took- 47, there goes $3 in 3 minutes, or half that time if your quick on the keyboard.

And even then! Well, you’re limited aren’t you? Limited to the Gemini model, or OpenAI, or whoever, and you see the limits of any one model in the article as well. Or you plonk down for a mediocre GPU with some slight VRAM headroom and choose from dozens of models, countless Lora, control nets, and other options, infinitely flexible in painting and outpainting. Ahead of that you’ll need to budget at least a dozen hours to learn local genai tools, comfyui or others. Then, for under a $1 dollar in electricity, you can can queue up a dozen ideas overnight and get 1,000 variations on each of them handed to you in the morning to quickly triage over coffee and email catchup.

It’s not a one size fits all market though, and most professionals are likely finding they want both: A low-cost, high-control, high precision sandbox that isn’t as fast or scalable as the api, and the api for when fast and scalable is what you need.

Re: Generative AI Image Editing Showdown

#79
post #41

Earlier quoted context omitted.

GPUs are needed for plenty of reasons. I assume plenty have a decent dGPU, even on laptops.

I have a 4080 RTX and Kontext runs great at fp8. I run several other models besides. If you want to get at all good at this, you need tons of throwaway generations and fast iteration and an API quickly becomes pricier than a GPU.

Precisely. Even inflated if the inflated 16,000 api calls was accurate for how much the cost of mediocre GPU would get you, that’s not an endless store of api calls. I’m also on a 4080 for lighter loads, and even just writing benchmarks, exploring attention mechanisms, token salience, etc, without image gen being my specific purpose I may trash half a thousand generations from output every few days. More if I count the stuff that never made it that far too.

Re: Generative AI Image Editing Showdown

#80
This is so much more useful than synthetic benchmarks. The most important column here isn't pass/fail, it's attempts. In production a model that gets it right in 2 attempts is 10x more valuable than one that needs 20 iterations of prompt engineering. It's a direct measure of cost and predictability.

Seedream 4 won on points, but Gemini seems more steerable and required less fighting on many of the tasks

Post reply on HN