Live data from Hacker News

Gemini 2.5 Flash Image

developers.googleblog.com

281–290 of 504 posts

Re: Gemini 2.5 Flash Image

#281
post #112
post #60

This is the gpt 4 moment for image editing models. Nano banana aka gemini 2.5 flash is insanely good. It made a 171 elo point jump in lmarena! Just search nano banana on Twitter to see the crazy results. An example. https://x.com/D_studioproject/status/1958019251178267111

Alarming hands on the third one: it can't decide which way they're facing. But Gemini didn't introduce that, it's there in the base image.

Yes, the base image's hands are creepy.

Re: Gemini 2.5 Flash Image

#282

Earlier quoted context omitted.

[flagged]

Please pardon me since I don't know if this is satirical or not. I'd wish if you could clarify it. Because if this is real, then the world is cooked if not, then the fact that I think that It might be real but the only reason I believe its a joke is because you are on hackernews so I think that either you are joking or the tech has gotten so convincing that even people on hackernews (which I hold to a fair standard)…

Not satire. He made a big speech about rewarding those who invested early in tech to move humanity forward and the benefits of the blockchain. It was extremely convincing. Three college grads and a medical doctor were all convinced.

Re: Gemini 2.5 Flash Image

#283

I digitised our family photos but a lot of them were damaged (shifted colours, spills, fingerprints on film, spots) that are difficult to correct for so many images. I've been waiting for image gen to catch up enough to be able to repair them all in bulk without changing details, especially faces. This looks very good at restoring images without altering details or adding them where they are missing, so it might fina…

I don't really understand the point of this usecase. Like, can't you also imagine what the photos might look like without the damage? Same with AI upscaling in phone cameras... if I want a hypothetical idea of what something in the distance might look like, I can just... imagine it? I think we will eventually have AI based tools that are just doing what a skilled human user would do in Photoshop, via tool-use. This w…

Read up on aphantasia.

Re: Gemini 2.5 Flash Image

#284
post #236

Earlier quoted context omitted.

I am flabbergasted that you both get scammed. I would understand if this was two years ago, but now? Do people really not know about these scams? I can already see down votes coming for victim blaming, but this is to me really shocking. Notice that there isn't "tell hn: don't get scammed by deep fake crypto Elon" because people who usually posts also consider this general knowledge. That's why it's so effective I gue…

Hey it takes courage to admit to it. That’s admirable.

I am so deeply ashamed.

Re: Gemini 2.5 Flash Image

#285

Like most image generators, it didn’t pass the piano keyboard test. (Black keys are wrong.) https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%...

Are their models that have vector space that includes ideas, not just words/media but not entirely corporeal aspects? So when generating a video of someone playing a keyboard the model would incorporate the idea of repeating groups of 8 tones, which is a fixed ideational aspect which might not be strongly represented in words adjacent to "piano". It seems like models need help with knowing what should be static, or h…

How would you encode those ideas?

Re: Gemini 2.5 Flash Image

#286

Earlier quoted context omitted.

[flagged]

Would you consider writing a blog post about this experience? I'm incredibly interested in learning more details about how this unfolded.

Yeah, here it is, along with screenshots. https://docs.google.com/document/d/1lRbApgKT4U95zN0AYsPQqsLR...

Re: Gemini 2.5 Flash Image

#287
post #279

I've updated the GenAI Image comparison site (which focuses heavily on strict text-to-image prompt adherence) to reflect the new Google Gemini 2.5 Flash model (aka nano-banana). https://genai-showdown.specr.net This model gets 8 of the 12 prompts correct and easily comes within striking distance of the best-in-class models Imagen and gpt-image-1 and is a significant upgrade over the old Gemini Flash 2.0 model. The re…

Why do Hunyuan, OpenAI 4o and Gwen get a pass for the octopus test? They don't cover "each tentacle", just some. And midjourney covers 9 of 8 arms with sock puppets.

Good point. I probably need to adjust the success pass ratios to be a bit stricter, especially as the models get better.

> midjourney covers 9 of 8 arms with sock puppets.

Midjourney is shown as a fail so I'm not sure what your point is. And those don't even look remotely close to sock puppets, they resemble stockings at best.

Re: Gemini 2.5 Flash Image

#288

Earlier quoted context omitted.

[flagged]

I don't mean to be rude, but this sounds like natural selection doing its work.

I'm pretty successful with an above average IQ. It was very convincing, along with three other college grads (one a medical doctor).

Re: Gemini 2.5 Flash Image

#289
I'm really waiting for a Pro sized Gemini model with image output.

I experimented heavily with 2.0 for a site I work on, but it never left preview and it had some gaps that were clearly due to being a small model (like lacking world knowledge, struggling with repetition, missing nuance in instructions, etc.)

2.5 Flash/nano-banana is a major step up but still has small model gaps peeking through. It still gets to "locked in" states where it's just repeating itself, which is a similar failure mode of small models for creative writing tasks.

A 2.5 Pro would likely close those gaps and definitively beat gpt-image-1

Re: Gemini 2.5 Flash Image

#290

Earlier quoted context omitted.

What is the piano keyboard test? Your link requires granting AI Studio access to Google Drive, which I do not want to do.

Just ask it to generate a correct piano keyboard. It's something the current gen of image generator AIs fail at.

Do most humans pass?
Post reply on HN