Live data from Hacker News

Gemini 2.5 Flash Image

developers.googleblog.com

271–280 of 504 posts

Re: Gemini 2.5 Flash Image

#271

Earlier quoted context omitted.

I think the comment is a joke. Their bio is satirical at least :)

I'm pretty sure the comment wasn't a joke? I saw the stream last week, it was very impressive use of AI, I didn't realize it was AI until he started talking about doubling crypto. What about the bio is satirical? I'm pretty sure that's sincere too.

User has edited their bio now :)

Re: Gemini 2.5 Flash Image

#272

I digitised our family photos but a lot of them were damaged (shifted colours, spills, fingerprints on film, spots) that are difficult to correct for so many images. I've been waiting for image gen to catch up enough to be able to repair them all in bulk without changing details, especially faces. This looks very good at restoring images without altering details or adding them where they are missing, so it might fina…

I don't really understand the point of this usecase. Like, can't you also imagine what the photos might look like without the damage? Same with AI upscaling in phone cameras... if I want a hypothetical idea of what something in the distance might look like, I can just... imagine it? I think we will eventually have AI based tools that are just doing what a skilled human user would do in Photoshop, via tool-use. This w…

Not everyone has a great imagination.

Re: Gemini 2.5 Flash Image

#273

Earlier quoted context omitted.

I've been testing it against Flux Pro Kontext for several weeks. I would say it beats Flux in a majority of tests, but Flux still surprises from time-to-time. Banana definitely isn't the best 100% of the time -- it falls a bit short of that. Evolution, not revolution.

Agreed. I find myself alternating between Qwen Image Edit 20B, Kontext, and now Flash 2.5 depending on the situation and style. And of course, Flash isn't open-weights, so if you need more control / less censorship then you're SOL.

Has there been a sufficient indication to conclude these weights will not (now or ever) be released?

Re: Gemini 2.5 Flash Image

#274

At $0.02 per image, it's prohibitively expensive for many use-cases. For comparison, the cheapest Flux model (Schnell) is $0.003 per image.

How many images do you need? What are the use-cases that need a bunch of artificial yet photoreal images produced or altered without human supervision?

Re: Gemini 2.5 Flash Image

#275

I am glad that I never decided to become a photoshop pro. I always contemplated about it, seemed attractive for a while, but glad that I decided against it. RIP r/photoshopbattles. It was in the endless list of new shiny 'skills' that feels good to have. Now I can use nano-banana instead. Other models will soon follow, I am sure.

it's still a useful skill to know photoshop. AI images can be great but you are almost always going to want to A. create the base composition yourself B. clean up artifacts in the AI generation and C. layer AI compositions into a final work.

Re: Gemini 2.5 Flash Image

#276

I tried to reproduce the fork/spaghetti example and the fashion bubble example, and neither looks anything like what they present. The outputs are very consistent, too. I am copying/pasting the images out of the advertisement page so they may be lower resolution than the original inputs, but otherwise I'm using the same prompts and getting a wildly different result. It does look like I'm using the new model, though.…

Wildly different and subjectively less "presentable", to be clear. The fashion bubble just generates a vague bubble shape with the subject inside it instead of the"subject flying through the sky inside a bubble" presented on the site. The other case just adds the fork to the bowl of spaghetti. Both are reproducible.

Arguably they follow the prompt better than what Google is showing off, but at the same time look less impressive.

Re: Gemini 2.5 Flash Image

#277

Earlier quoted context omitted.

Would you consider writing a blog post about this experience? I'm incredibly interested in learning more details about how this unfolded.

I think the comment is a joke. Their bio is satirical at least :)

Their bio mentions their actual job and one project that is verifiably real. I think that the elements that seem satirical are real projects they're working on.

Re: Gemini 2.5 Flash Image

#278

I tried to reproduce the fork/spaghetti example and the fashion bubble example, and neither looks anything like what they present. The outputs are very consistent, too. I am copying/pasting the images out of the advertisement page so they may be lower resolution than the original inputs, but otherwise I'm using the same prompts and getting a wildly different result. It does look like I'm using the new model, though.…

The output consistency is interesting. I just went through half a dozen generations of my standard image model challenge, (to date I have yet to see a model that can render piano keyboard octaves correctly, and Gemini 2.5 Flash Image is no different in that regard), and as best I can tell, there are no changes at all between successive attempts: https://g.co/gemini/share/a0e1e264b5e9 This is in stark contrast to Chat…

Flash 2.0 Image had the same issue: it does better than gpt-image for maintaining consistency in edits, but that also introduces a gap where sometimes it gets "locked in" on a particular reference image and will struggle to make changes to it.

In some cases you'll pass in multiple images + a prompt and get back something that's almost visually indistinguishable from just one of the images and nothing from the prompt.

Re: Gemini 2.5 Flash Image

#279

I've updated the GenAI Image comparison site (which focuses heavily on strict text-to-image prompt adherence) to reflect the new Google Gemini 2.5 Flash model (aka nano-banana). https://genai-showdown.specr.net This model gets 8 of the 12 prompts correct and easily comes within striking distance of the best-in-class models Imagen and gpt-image-1 and is a significant upgrade over the old Gemini Flash 2.0 model. The re…

Why do Hunyuan, OpenAI 4o and Gwen get a pass for the octopus test? They don't cover "each tentacle", just some. And midjourney covers 9 of 8 arms with sock puppets.

Re: Gemini 2.5 Flash Image

#280

Earlier quoted context omitted.

I'm pretty sure the comment wasn't a joke? I saw the stream last week, it was very impressive use of AI, I didn't realize it was AI until he started talking about doubling crypto. What about the bio is satirical? I'm pretty sure that's sincere too.

User has edited their bio now :)

I didn't edit my bio. My projects are not satire. I'm just less ashamed than most, so I work on more "exciting" projects. I've worked extensively with generative AI, including video, myself. It was just that convincing to me in the moment. My regret knows no bounds. Luckily I earn enough this doesn't devastate me, but I really could have done some good with that money.
Post reply on HN