Live data from Hacker News

Gemini 2.5 Flash Image

developers.googleblog.com

371–380 of 504 posts

Re: Gemini 2.5 Flash Image

#371

I've updated the GenAI Image comparison site (which focuses heavily on strict text-to-image prompt adherence) to reflect the new Google Gemini 2.5 Flash model (aka nano-banana). https://genai-showdown.specr.net This model gets 8 of the 12 prompts correct and easily comes within striking distance of the best-in-class models Imagen and gpt-image-1 and is a significant upgrade over the old Gemini Flash 2.0 model. The re…

What's interesting is that Imagen 4 and Gemini 2.5 Flash Image look suspiciously similar in several of these tests cases. Maybe Gemini 2.5 Flash first calls Imagen in the background to get a detailed baseline image (diffusion models are good at this) and then Gemini edits the resulting image for better prompt adherence.

Re: Gemini 2.5 Flash Image

#372

Unfortunately, it suffers from the same safetyism than other many releases. Half of the prompts get rejected. How can you have character consistency if the model is forbidden from editing any human. And most of my photo editing involves humans, so basically this is just a useless product. I get that Google doesn't want to be responsible for deep fake advances, but that seems inevitable, so this is just slightly delay…

I have an old photo of my girlfriend with her cousin when they were young, wearing Christmas dresses in front of the tree, not long before they were separated to other sides of the world for decades now. The photo is itself low quality on top of the photo itself being physically beat up. So far no model is willing to clean it up :/

There are reddit communities (I admittedly don't remember which, but could probably be found from a simple search) where people will offer their photo editing skills to touch up the photo, often for free. Could be worth trying a real human if the robots are going full HAL 9000 and telling you they can't do it.

Re: Gemini 2.5 Flash Image

#373

Earlier quoted context omitted.

What tools did you use to make those videos from the PG image?

I used a bunch of models in conjunction: - Midjourney (background) - Qwen Image (restyle PG) - Gemini 2.5 Flash (editing in PG) - Gemini 2.5 Flash (adding YC logo) - Kling Pro (animation) I didn't spend too much time correcting mistakes. I used a desktop model aggregation and canvas tool that I wrote [1] to iterate and structure the work. I'll be open sourcing it soon. [1] https://getartcraft.com

The app looks interesting, but I think it needs some documentation. I think I generated something? Maybe? I saw a spinny thing for awhile, but then nothing.

I couldn't get the 3d thing to do much. I had assets in the scene but I couldn't for the life of me figure out how to use the move, rotate or scale tools. And the people just had their arms pointing outward. Are you supposed to pose them somehow? Maybe I'm supposed to ask the AI to pose them?

Inpainting I couldn't figure out either... It's for drawing things into an existing image (I think?) but it doesn't seem to do anything other than show a spinny thing for awhile...

I didn't test the video tool because I don't have a midjourney account.

Re: Gemini 2.5 Flash Image

#374

All images created or edited with Gemini 2.5 Flash Image will include an invisible SynthID digital watermark, so they can be identified as AI-generated or edited. Obviously I understand what is the purpose and the good intention, but I think sad to see that we are not not anymore responsible adults but big corps deciding for us what we can and what we cannot do. Snitching on your back.

I'm generally against "if you have not thing to fear you have nothing to hide" arguments but I'm curious what your argument is here for why it would be a problem that AI generated and edited images can be recognized as such.

Edit: I should probably say for full transparency that I am strongly FOR watermarks for AI imagery

Re: Gemini 2.5 Flash Image

#375

All images created or edited with Gemini 2.5 Flash Image will include an invisible SynthID digital watermark, so they can be identified as AI-generated or edited. Obviously I understand what is the purpose and the good intention, but I think sad to see that we are not not anymore responsible adults but big corps deciding for us what we can and what we cannot do. Snitching on your back.

Also, it's not really your image. Like if an artist puts a watermark on a commissioned piece it's not really a good argument that the artist is "snitching" by saying the art was done by them and not you trying to pass it off as your own...

I don't know if that's the argument you're trying to make, but I think it's worth considering

Re: Gemini 2.5 Flash Image

#376
post #261

Earlier quoted context omitted.

What are "the arenas"?

Blind rating battlegrounds, one is https://lmarena.ai/ (first google result)

I don't quite get what this is? I asked the AI on the site "What is imarena.ai?" and it just gave some hallucinated answer that made no sense.

Re: Gemini 2.5 Flash Image

#377

Earlier quoted context omitted.

did you see the generated pic demis posted on X? it looks like slop from 2 years ago. https://x.com/demishassabis/status/1960355658059891018

I've tested it on Google AI Studio since it's available to me (which is just a few hours so take it with a grain of salt). The prompt comprehension is uncannily good. My test is going to https://unsplash.com/s/photos/random and pick two random images, send them both and "integrate the subject from the second image into the first image" as the prompt. I think Gemini 2.5 is doing far better than ChatGPT (admittedly Cha…

> FluxKontext

Flux Kontext is an editing model, but the set of things it can do is incredibly limited. The style of prompting is very bare bones. Qwen (Alibaba) and SeedEdit (ByteDance) are a little better, but they themselves are nowhere near as smart as Gemini 2.5 Flash or gpt-image-1.

Gemini 2.5 Flash and gpt-image-1 are in a class of their own. Very powerful instructive image editing with the ability to understand multiple reference images.

> Edit: Honestly it might not be the 'gpt4 moment." It's better at combining multiple images, but now I don't think it's better at understanding elaborated text prompt than ChatGPT.

Both gpt-image-1 and Gemini 2.5 Flash feel like "Comfy UI in a prompt", but they're still nascent capabilities that get a lot wrong.

When we get a gpt-image-1 with Midjourney aesthetics, better adherence and latency, then we'll have our "GPT 4" moment. It's coming, but we're not there yet.

They need to learn more image editing tricks.

Re: Gemini 2.5 Flash Image

#378
post #60

This is the gpt 4 moment for image editing models. Nano banana aka gemini 2.5 flash is insanely good. It made a 171 elo point jump in lmarena! Just search nano banana on Twitter to see the crazy results. An example. https://x.com/D_studioproject/status/1958019251178267111

The fingernails on one of them. Ohhh nooo

Re: Gemini 2.5 Flash Image

#379

All images created or edited with Gemini 2.5 Flash Image will include an invisible SynthID digital watermark, so they can be identified as AI-generated or edited. Obviously I understand what is the purpose and the good intention, but I think sad to see that we are not not anymore responsible adults but big corps deciding for us what we can and what we cannot do. Snitching on your back.

dont worry, you can just screenshot the image to get rid of the watermark

That's not correct. The watermark is robust to screenshots, file format changes (saving as jpeg/png) and at least light transformation (cropping, saturation level adjustment, etc).
Post reply on HN