Live data from Hacker News

Generative AI Image Editing Showdown

genai-showdown.specr.net

21–30 of 83 posts

Re: Generative AI Image Editing Showdown

#21
It is fun being one of the elderly who set their standards back in distant 2022. All these demos look incredible compared to SD1, 2 & 3. We've entered a very different era where the models seem to actually understand both the prompt and the image instead of throwing paint at the wall in a statistically interesting manner.

I think this was fairly predictable, but as engineering improvements keep happening and the prompt adherence rate tightens up we're enjoying a wild era of unleashed creativity.

Re: Generative AI Image Editing Showdown

#24

I wonder how much longer those annoying stock photo database will continue. They are great for press photography and such. But stock pics of people in offices for a website are nothing, I would buy a min 3 month subscription for anymore

As generative AI eats away at the high royalty, restrictive license, consent evading, stereotype reinforcing business model of stock photo companies, it will be a challenge to resist the schadenfreude.

Re: Generative AI Image Editing Showdown

#25

I do not use ai image generating much lately. It seemed like there was a burst of activity a year and half ago with self hosted models and using some localhost web guis. But now it seems like it is moving more and more to online hosted models. Still, to my eye, ai generated images still feel a bit off when doing with real world photographs. George's hair, for example, looks over the top, or brushed on. The tree added…

> But now it seems like it is moving more and more to online hosted models. It's mostly because image model size and required compute for both training and inference have grown faster than self-hosted compute capability for hobbyists. Sure, you can run Flux Kontext locally, but if you have to use a heavily quantized model and wait forever for the generation to actually run, the economics are harder to justify. That's…

For 99% of my use cases I’ll just use ChatGPT or Gemini due to convenience. But if you want something with a specific style, Flux LoRAs are much better, in which case I’ll boot up the old 4090.

The economics 1000% do not justify me owning a GPU to do this. I just happen to own one.

Re: Generative AI Image Editing Showdown

#26

Everyone is sleeping on Gemini 2.5 Flash Image / Nano Banana. As shown in the OP, it's substantially more powerful than most other models while at the same price-per-image, and due to its text encoder it can handle significantly larger and more nuanced prompts to get exactly what you want. I open-sourced a Python package for generating from it with examples ( https://github.com/minimaxir/gemimg ) and am currently wor…

[deleted]

Re: Generative AI Image Editing Showdown

#27

Everyone is sleeping on Gemini 2.5 Flash Image / Nano Banana. As shown in the OP, it's substantially more powerful than most other models while at the same price-per-image, and due to its text encoder it can handle significantly larger and more nuanced prompts to get exactly what you want. I open-sourced a Python package for generating from it with examples ( https://github.com/minimaxir/gemimg ) and am currently wor…

I don't think people are really sleeping on it - nano-banana more or less went viral when it first came out. I'd argue that aside from the capabilities built into ChatGPT (with the Ghibli craze and whatnot) craze it's the best known image editing model.

It's a weird situation where the Gemini mobile app hit #2 on the App Stores because of free Nano Banana, but no one ever talks about it and most disclosed image generations I've seen are still ChatGPT.

Re: Generative AI Image Editing Showdown

#28

I'm pretty sure that "replace the homeless man with a park bench" image was a reference to some TV show making a gentrification joke, but I can't put my finger on it. Anyone recall?

Simpsons, the Frank Scorpio episode. The advertisement for the company town shows a beggar slowly fading out and being replaced by a mailbox.

Re: Generative AI Image Editing Showdown

#29

Everyone is sleeping on Gemini 2.5 Flash Image / Nano Banana. As shown in the OP, it's substantially more powerful than most other models while at the same price-per-image, and due to its text encoder it can handle significantly larger and more nuanced prompts to get exactly what you want. I open-sourced a Python package for generating from it with examples ( https://github.com/minimaxir/gemimg ) and am currently wor…

No one is sleeping on nano-banana/Gemini Flash, it's highly over-tuned for editing vs novel generation and maxes out at a pretty low resolution.

Seedream 4.0 is somewhat slept on for being 4k at the same cost as nano-banana. It's not as great at perfect 1:1 edits, but it's aesthetics are much better and it's significantly more reliable in production for me.

Models with LLM backbones/omni-modal models are not rare anymore, even Qwen Image Edit is out there for open-weights.

Re: Generative AI Image Editing Showdown

#30

I'm pretty sure that "replace the homeless man with a park bench" image was a reference to some TV show making a gentrification joke, but I can't put my finger on it. Anyone recall?

Yeah, I couldn't help myself on that one! It's a reference to the Cypress Creek promotional video from the Simpsons.

https://www.youtube.com/watch?v=foU9W7AkKSY

Post reply on HN