Live data from Hacker News

Gemini 2.5 Flash Image

developers.googleblog.com

181–190 of 504 posts

Re: Gemini 2.5 Flash Image

#181

I tried to reproduce the fork/spaghetti example and the fashion bubble example, and neither looks anything like what they present. The outputs are very consistent, too. I am copying/pasting the images out of the advertisement page so they may be lower resolution than the original inputs, but otherwise I'm using the same prompts and getting a wildly different result. It does look like I'm using the new model, though.…

The output consistency is interesting. I just went through half a dozen generations of my standard image model challenge, (to date I have yet to see a model that can render piano keyboard octaves correctly, and Gemini 2.5 Flash Image is no different in that regard), and as best I can tell, there are no changes at all between successive attempts: https://g.co/gemini/share/a0e1e264b5e9

This is in stark contrast to ChatGPT, where an edit prompt typically yields both requested and unrequested changes to the image; here it seems to be neither.

Re: Gemini 2.5 Flash Image

#183
post #87

Earlier quoted context omitted.

Hope it works well for you! In my eyes, one specific example they show (“Prompt: Restore photo”) deeply AI-ifies the woman’s face. Sure it’ll improve over time of course.

I tried a dozen or so images. For some it definitely failed (altering details, leaving damage behind, needing a second attempt to get a better result) but on others it did great. With a human in the loop approving the AI version or marking it for manual correction I think it would save a lot of time. This is the first image I tried: https://i.imgur.com/MXgthty.jpeg (before) https://i.imgur.com/Y5lGcnx.png (after) Sur…

Being pragmatic, the after is a good restoration. There is nothing really lost (except some sharpness that could be put back). The main failing of AI is on faces because our brains are so hardwired to see any changes or weirdness. This is the sort of image that is perfect for AI because the subject's face is already occluded.

Re: Gemini 2.5 Flash Image

#184

Like most image generators, it didn’t pass the piano keyboard test. (Black keys are wrong.) https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%...

Interesting! I feel like that's maybe similar to the business of being able to correctly generate images of text— it looks like the idea of a keyboard to a non-musician, but is immediately wrong to someone who is actually familiar with it at all.

I wonder if the bot is forced to generate something new— certainly for a prompt like that it would be acceptable to just pick the first result off a google image search and be like "there, there's your picture of a piano keyboard".

Re: Gemini 2.5 Flash Image

#185

Like most image generators, it didn’t pass the piano keyboard test. (Black keys are wrong.) https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%...

What is the piano keyboard test? Your link requires granting AI Studio access to Google Drive, which I do not want to do.

Re: Gemini 2.5 Flash Image

#186
post #73

Earlier quoted context omitted.

I have been testing google's SynthID for images and while it isn't perfect, it is very good, insofar that I felt some relief from that same creeping dread over what these images will do to perceived reality. It survives a lot of transformation like compression, cropping, and resizing. It even survives over alterations like color filtering and overpainting.

facebook isn't going to implement detection though. Many (if not most) of the viral pictures are AI-generated. and facebook is incentivized to let their users get fooled to generate endless scrolling

They already did. Certainly on the backend. For a while they were surfacing it, but I think it's gone again. But Meta is definitely onto this.

Re: Gemini 2.5 Flash Image

#187
post #176

Are men not attractive? Or perhaps for Google, this blog is a targeted content? But who is it targeting? I would like to see the reasoning behind using all women images (at the least the top/first ones) to show off the model capabilities. I have noticed this trend in the image manipulation business a lot.

The average man finds the average woman more attractive than the average woman finds the average man. Replace attractive with (eye-catching/attention-grabbing/motivating/retention-boosting).

Oh, in that case, it makes sense. Also, I think men/women consume different kind of media and this is one of those "men dominated" corner of the internet. I also think due to trainig data bias - there could be some difference in quality with different subjects. So, they might be showing off their best of best.

Re: Gemini 2.5 Flash Image

#188

Super cheap generation but expensive image upload, do I read that right? https://openrouter.ai/google/gemini-2.5-flash-image-preview

Not sure. If the Flash image output is $30/M [1] then that's pretty similar to gpt-image-1 costs. So a faster and better model perhaps but not really cheaper?

[1] https://developers.googleblog.com/en/introducing-gemini-2-5-...

Re: Gemini 2.5 Flash Image

#190

Earlier quoted context omitted.

What a weird rejection. You have to scroll pretty far in the article to see an example output that doesn't have a realistic depiction of a person.

It's unfortunate they can't just explain the real reason they don't want to generate the image: "Unfortunately I'm not able to generate images that might cause bad PR for Alphabet(tm) or subsidiaries. Is there anything else I can generate for you?"

If you want that kind of thing, Qwen3 delivers:

https://www.reddit.com/r/LocalLLaMA/comments/1mx1pkt/qwen3_m...

Post reply on HN