Live data from Hacker News

Gemini 2.5 Flash Image

developers.googleblog.com

411–420 of 504 posts

Re: Gemini 2.5 Flash Image

#411

There is one thing Gemini 2.5 Flash Image can do that no other edit model can do: incorporate multiple images simultaneously without shenanigans due to its multimodality, e.g. for Flux Kontext, if you want to "put the person in the first image into the second image", you have to concatenate them pre-VAE which can be unwieldly, but this model doesn't have that issue. You can even incorporate more than two images, but…

I very much enjoy this feature. My next door neighbor is on vacation, and I'm feeding his fish for him. I took a picture of the fish tank and asked Gemini to put the fish tank at various local tourist attractions in my city, as if we're going on day trips. I send him one photo a day and he's been loving it. Just a fun little thing to put a smile on his face (and mine).

Fun fact - I trained a lora on our almost-toddler at the time on SDXL and generated images of her doing dangerous things to send to my wife the first day she had a trip away from us.

It was all fun and games until the little shit crawled out of our doggy door for the first and only time when I was going to the bathroom. As I was looking for her I got a notification we were in a tornado warning.

Luckily the dog knew where she had gone and led me to her, having crawled down our (3 step) deck, across our yard, and was standing looking up at the angry clouds.

Re: Gemini 2.5 Flash Image

#412

Earlier quoted context omitted.

Vibe coding might not be real, but vibe graphics design certainly is. https://imgur.com/a/internet-DWzJ26B Anyone can make images and video now.

Are those oil derricks, or wind turbines? Who cares! Graphic design is easy now!

[deleted]

Re: Gemini 2.5 Flash Image

#413

Earlier quoted context omitted.

If you compare to the amount of effort required in Photoshop to achieve the same results, still a vast improvement

Vibe coding might not be real, but vibe graphics design certainly is. https://imgur.com/a/internet-DWzJ26B Anyone can make images and video now.

[deleted]

Re: Gemini 2.5 Flash Image

#414

Unfortunately, it suffers from the same safetyism than other many releases. Half of the prompts get rejected. How can you have character consistency if the model is forbidden from editing any human. And most of my photo editing involves humans, so basically this is just a useless product. I get that Google doesn't want to be responsible for deep fake advances, but that seems inevitable, so this is just slightly delay…

I have an old photo of my girlfriend with her cousin when they were young, wearing Christmas dresses in front of the tree, not long before they were separated to other sides of the world for decades now. The photo is itself low quality on top of the photo itself being physically beat up. So far no model is willing to clean it up :/

If you have a decent GPU Qwen Edit can probably do it and certainly won’t refuse.

Keep in mind no editing model is magic and if the pixels just aren’t there for their faces it’s essentially going to be making stuff up.

Re: Gemini 2.5 Flash Image

#415

Is it capable of generating white males? Given that this has been a serious problem with Google models, I would guess it would have been a good thing to at least add one such example in the marketing material. If the marketing material is to be believed, the model really prefers black females.

> Given that this has been a serious problem with Google models

It hasn't been though, has it?

At one point one of their earliest image gen models had a prompting problem: they tried to have the LLM doing prompt expansion avoid always generating white people, since they realized white people were significantly overrepresented in their training data compared to their proportion of the population.

Unfortunately that prompt expansion would sometimes clash with cases where there was a specific race required for historical accuracy.

AFAIK they fixed that ages ago and it stopped being an issue.

Re: Gemini 2.5 Flash Image

#416
post #395
post #226

this is amazing. I just wish models would have more non-textual controls. I don't want to TYPE my instructions. We need a better UI for editing images with AI.

You want to use a mouse and keyboard and learn 20 buttons like it's 1990?

I don't want to converse with a 4 year old with the world's photo libraries at its disposal. I spent 10 minutes trying to convince the model to add a watch to a person's left arm instead of the right arm and it would not do it, it apparently could not get the idea on this particular image. If I had a drawing tool I could circle where I wanted it and say 'THERE, stupid'. Next year when we have AGI all of this will be moot of course, but for now Photoshop isn't going away.

Re: Gemini 2.5 Flash Image

#417
post #236

Earlier quoted context omitted.

I am flabbergasted that you both get scammed. I would understand if this was two years ago, but now? Do people really not know about these scams? I can already see down votes coming for victim blaming, but this is to me really shocking. Notice that there isn't "tell hn: don't get scammed by deep fake crypto Elon" because people who usually posts also consider this general knowledge. That's why it's so effective I gue…

Hey it takes courage to admit to it. That’s admirable.

Yes, thanks OP for sharing. I check HN front page mostly everyday and had no clue such sophisticated scams existed (I pretty much don’t use social media).

It’s easy to think “eh, it will never happen to me” but hindsight is 20/20. I impulse-donated to things like Wikipedia in the past and I’m susceptible to FOMO as most people.

Re: Gemini 2.5 Flash Image

#418
post #60

This is the gpt 4 moment for image editing models. Nano banana aka gemini 2.5 flash is insanely good. It made a 171 elo point jump in lmarena! Just search nano banana on Twitter to see the crazy results. An example. https://x.com/D_studioproject/status/1958019251178267111

I've been testing it for several weeks. It can produce results that are truly epic, but it's still a case of rerolling the prompt a dozen times to get an image you can use. It's not God. It's definitely an enormous step though, and totally SOTA.

The model seems good, but it seems to have huge issues in doing garbage most of times lol.

Still needs more RLHF tuning I guess? As the previous version was even worse.

Re: Gemini 2.5 Flash Image

#419

Earlier quoted context omitted.

I don't quite get what this is? I asked the AI on the site "What is imarena.ai?" and it just gave some hallucinated answer that made no sense.

People vote on the performance of AI, generating ranking boards.

Ah, that was the missing piece of information! Thanks!

Re: Gemini 2.5 Flash Image

#420

Earlier quoted context omitted.

I've been testing it against Flux Pro Kontext for several weeks. I would say it beats Flux in a majority of tests, but Flux still surprises from time-to-time. Banana definitely isn't the best 100% of the time -- it falls a bit short of that. Evolution, not revolution.

Agreed. I find myself alternating between Qwen Image Edit 20B, Kontext, and now Flash 2.5 depending on the situation and style. And of course, Flash isn't open-weights, so if you need more control / less censorship then you're SOL.

It’s good but holy shit is it censored! Try generating any kind of scene on a beach…
Post reply on HN