Live data from Hacker News

Gemini 2.5 Flash Image

developers.googleblog.com

261–270 of 504 posts

Re: Gemini 2.5 Flash Image

#261

Earlier quoted context omitted.

Why is it called nano banana?

Before a model is announced, they use codenames on the arenas. If you look online, you can see people posting about new secret models and people trying to guess whose model it is.

What are "the arenas"?

Re: Gemini 2.5 Flash Image

#262
post #60

This is the gpt 4 moment for image editing models. Nano banana aka gemini 2.5 flash is insanely good. It made a 171 elo point jump in lmarena! Just search nano banana on Twitter to see the crazy results. An example. https://x.com/D_studioproject/status/1958019251178267111

I wonder how the creative workflow looks like when this kind of models are natively integrated into digital image tools. Imagine fine-grained controls on each layer and their composition with the semantic understanding on the full picture.

Re: Gemini 2.5 Flash Image

#263
post #60

This is the gpt 4 moment for image editing models. Nano banana aka gemini 2.5 flash is insanely good. It made a 171 elo point jump in lmarena! Just search nano banana on Twitter to see the crazy results. An example. https://x.com/D_studioproject/status/1958019251178267111

I've been testing it for several weeks. It can produce results that are truly epic, but it's still a case of rerolling the prompt a dozen times to get an image you can use. It's not God. It's definitely an enormous step though, and totally SOTA.

How did you get early access? Thanks.

Re: Gemini 2.5 Flash Image

#265
post #261

Earlier quoted context omitted.

Before a model is announced, they use codenames on the arenas. If you look online, you can see people posting about new secret models and people trying to guess whose model it is.

What are "the arenas"?

Blind rating battlegrounds, one is https://lmarena.ai/ (first google result)

Re: Gemini 2.5 Flash Image

#266

I am glad that I never decided to become a photoshop pro. I always contemplated about it, seemed attractive for a while, but glad that I decided against it. RIP r/photoshopbattles. It was in the endless list of new shiny 'skills' that feels good to have. Now I can use nano-banana instead. Other models will soon follow, I am sure.

Interesting take. I'm a programmer, but learned Photoshop in the early 2000s and had a blast making and editing images for fun. Sure, the generative models today can do a far better job than anything I could come up with, but that doesn't detract from the experience and skills I picked up over the years. If anything, knowing Photoshop (I use Affinity Designer/Photo these days) is actually incredibly useful to finesse…

[deleted]

Re: Gemini 2.5 Flash Image

#267

Earlier quoted context omitted.

I've been testing it for several weeks. It can produce results that are truly epic, but it's still a case of rerolling the prompt a dozen times to get an image you can use. It's not God. It's definitely an enormous step though, and totally SOTA.

Is it because the model is not good enough at following the prompt, or because the prompt is unclear? Something similar has been the case with text models. People write vague instructions and are dissatisfied when the model does not correctly guess their intentions. With image models it's even harder for model to guess it right without enough details.

Remember in image editing, the source image itself is a huge part of the prompt, and that's often the source of the ambiguity. The model may clearly understand your prompt to change the color of a shirt, but struggle to understand the boundaries of the shirt. I was just struggling to use AI to edit an image where the model really wanted the hat in the image to be the hair of the person wearing it. My guess for that bias is that it had just been trained on more faces without hats than with them on.

Re: Gemini 2.5 Flash Image

#268

Earlier quoted context omitted.

I've been testing it for several weeks. It can produce results that are truly epic, but it's still a case of rerolling the prompt a dozen times to get an image you can use. It's not God. It's definitely an enormous step though, and totally SOTA.

If you compare to the amount of effort required in Photoshop to achieve the same results, still a vast improvement

Vibe coding might not be real, but vibe graphics design certainly is.

https://imgur.com/a/internet-DWzJ26B

Anyone can make images and video now.

Re: Gemini 2.5 Flash Image

#270

Earlier quoted context omitted.

> This is the gpt 4 moment for image editing models. No it's not. We've had rich editing capabilities since gpt-image-1, this is just faster and looks better than the (endearingly? called) "piss filter". Flux Kontext, SeedEdit, and Qwen Edit are all also image editing models that are robustly capable. Qwen Edit especially. Flux Kontext and Qwen are also possible to fine tune and run locally. Qwen (and its video gen s…

I'm confused as well, I thought gpt-image could already do most of these things, but I guess the key difference is that gpt-image is not good for single point edits. In terms of "wow" factor it doesn't feel as big as gpt 3->4 though, since it sure _felt_ like models could already do this.

People really slept on gpt-image-1 and were too busy making Miyazaki/Ghibli images.

I feel like most of the people on HN are paying attention to LLMs and missing out on all the crazy stuff happening with images and videos.

LLMs might be a bubble, but images and video are not. We're going to have entire world simulation in a few years.

Post reply on HN