Live data from Hacker News

Gemini 2.5 Flash Image

developers.googleblog.com

211–220 of 504 posts

Re: Gemini 2.5 Flash Image

#211
post #60

This is the gpt 4 moment for image editing models. Nano banana aka gemini 2.5 flash is insanely good. It made a 171 elo point jump in lmarena! Just search nano banana on Twitter to see the crazy results. An example. https://x.com/D_studioproject/status/1958019251178267111

I've been testing it for several weeks. It can produce results that are truly epic, but it's still a case of rerolling the prompt a dozen times to get an image you can use. It's not God. It's definitely an enormous step though, and totally SOTA.

Is it because the model is not good enough at following the prompt, or because the prompt is unclear?

Something similar has been the case with text models. People write vague instructions and are dissatisfied when the model does not correctly guess their intentions. With image models it's even harder for model to guess it right without enough details.

Re: Gemini 2.5 Flash Image

#212

After the rugpull of Android, are we really going to trust Google with anything?

What does the first phrase even mean?

I think I should have used the word enshittification of Android. And, I need to brush up my writing, it's getting progressively worse.

Re: Gemini 2.5 Flash Image

#213

Earlier quoted context omitted.

All of the defects you have listed can be automatically fixed by using a film scanner with ICE and a software that automatically performs the scan and the restoration like Vuescan. Feeding hundreds (thousands?) of photos to an experimental proprietary cloud AI that will give you back subpar compressed pictures with who knows how many strange artifacts seems unnecessary

I scanned everything into 48-bit RAW and treat those as the originals, including the IR scan for ICE and a lower quality scan of the metadata. The problem is sharing them - important images I manually repair and export as JPEG which is time consuming (15-30 minutes per image, there are about 14000 total) so if its "generic family gathering picture #8228" I would rather let AI repair it, assuming it doesn't butcher fa…

this reminds me of a joke we used to tell as kids when there was a new Photoshop version coming out - "this one will remove the cow from the picture and we'll finally see what great-grandpa looked like!"

Re: Gemini 2.5 Flash Image

#214

Are men not attractive? Or perhaps for Google, this blog is a targeted content? But who is it targeting? I would like to see the reasoning behind using all women images (at the least the top/first ones) to show off the model capabilities. I have noticed this trend in the image manipulation business a lot.

Because tech is largely male dominated and has inherent sexism/patriarchy and images of women, especially conventionally attractive ones, has the perception of aiding sales.

Also women are seen as more cooperative and submissive, hence so many home assistants and AI being women's voices/femme coded.

Re: Gemini 2.5 Flash Image

#215
post #151

Earlier quoted context omitted.

As others have said, this is an image editing model. Editing models do not excel at aesthetic, but they can take your Midjourney image, adjust the composition, and make it perfect. These types of models are the Adobe killer.

Noted that! The editing capabilities are impressive. I was excited for image gen because of the API (Midjourney doesn't have it yet).

David Holz mentioned on Twitter that he was considering a Midjourney API. They're obviously providing it to Meta now, so it might become more broadly available after Midjourney becomes the default image gen for Meta products.

Midjourney wins on aesthetic for sure. Nothing else comes close. Midjourney images are just beautiful to behold.

David's ambition is to beat Google to building a world model you can play games in. He views the image and video business as a temporary intermediate to that end game.

Re: Gemini 2.5 Flash Image

#216

Like most image generators, it didn’t pass the piano keyboard test. (Black keys are wrong.) https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%...

The selling point of this model really seems to be it's consistency between generations rather than it's raw generating ability.

for instance:

https://aistudio.google.com/app/prompts/1gTG-D92MyzSKaKUeBu2...

Re: Gemini 2.5 Flash Image

#217
post #38

Very impressive. I have to say while I'm deeply impressed by these text to image models, there's a part of me that's also wary of their impact. Just look at the comments beneath the average Facebook post.

[flagged]

I got scammed similarly (although $10, because I tested first), because 1. it was on YouTube, on a channel called "SpaceX" with verified logo 2. with hundreds of thousands of viewers live 3. with a believable speech from Mr. Musk standing next to its rockets (and knowing his interest in cryptocurrencies).

This happened as I was genuinely searching for the actual live stream of SpaceX.

I am ashamed, even more so because I even posted the live stream link on Hacker News (!). Fortunately it was flagged early and I apologized personally to dang.

This was a terrible experience for me, on many levels. I never thought I would fall in such a trap, being very aware of the tech, reading about similar stories etc.

Re: Gemini 2.5 Flash Image

#218

Earlier quoted context omitted.

I've been testing it for several weeks. It can produce results that are truly epic, but it's still a case of rerolling the prompt a dozen times to get an image you can use. It's not God. It's definitely an enormous step though, and totally SOTA.

If you compare to the amount of effort required in Photoshop to achieve the same results, still a vast improvement

I work in Photoshop all day, and I 100% agree. Also, I just retried a task that wouldn't work last night on nano-banana and it worked first time on the released model, so I'm wondering if there were some changes to the released version?

Re: Gemini 2.5 Flash Image

#219

I am glad that I never decided to become a photoshop pro. I always contemplated about it, seemed attractive for a while, but glad that I decided against it. RIP r/photoshopbattles. It was in the endless list of new shiny 'skills' that feels good to have. Now I can use nano-banana instead. Other models will soon follow, I am sure.

Retouching is an art. To the pro, this is just another tool to increase efficiency. You pay them not just for knowing how to use Photoshop, but for exercising good judgement. That said, I imagine this will shrink the field, since fewer retouchers will be able to do the same work, unless the amount of work goes up commensurately. Will people get more retouching done if the price goes down? Not sure.

Re: Gemini 2.5 Flash Image

#220

I am glad that I never decided to become a photoshop pro. I always contemplated about it, seemed attractive for a while, but glad that I decided against it. RIP r/photoshopbattles. It was in the endless list of new shiny 'skills' that feels good to have. Now I can use nano-banana instead. Other models will soon follow, I am sure.

If you commented it a decade ago, I would say that at least you own the program and skills in case Google decides to turn off the lights or ask prohibitive price tag. Now you need to pay subscription for PS and maybe there would be some decent open weight model released.
Post reply on HN