Live data from Hacker News

Gemini 2.5 Flash Image

developers.googleblog.com

431–440 of 504 posts

Re: Gemini 2.5 Flash Image

#431

Half the time I ask Gemini to generate some image it claims it doesn't have the capability. And in general I've felt it's so hard to actually use the features Google announce? Like, a third of them is in one product, some in another which I can't use, and no idea what or where I should pay to get access. So confusing.

It's not in the Gemini app or site at all. You have to use AI Studio or another means. Yes, this is all very confusing on Google's part.

Hmm could the old models generate images before? Had they hooked up imagen or something? I can make images on the Gemini site.

Re: Gemini 2.5 Flash Image

#432

All images created or edited with Gemini 2.5 Flash Image will include an invisible SynthID digital watermark, so they can be identified as AI-generated or edited. Obviously I understand what is the purpose and the good intention, but I think sad to see that we are not not anymore responsible adults but big corps deciding for us what we can and what we cannot do. Snitching on your back.

I'm generally against "if you have not thing to fear you have nothing to hide" arguments but I'm curious what your argument is here for why it would be a problem that AI generated and edited images can be recognized as such. Edit: I should probably say for full transparency that I am strongly FOR watermarks for AI imagery

My problem is more the general idea that nowadays the tech is hostile to the user. Before when you paid for something, it was fully yours to use in a good or in a bad way.

Imagine for example, that in the future and with improved tech, manufacturer of knifes were to embed a gps chip in all knifes sold because it might be used for dangerous things.

The worse in all of that being that the big tech does it based on their own "moral" compass and not based on a legal requirement.

Regarding the watermark, that is also applying to generated text in theory, imagine that you ask ai to refactor a job application letter or a letter to your landlord, and that Google will snitch you with watermark that you used AI for that.

Re: Gemini 2.5 Flash Image

#433

Unfortunately, it suffers from the same safetyism than other many releases. Half of the prompts get rejected. How can you have character consistency if the model is forbidden from editing any human. And most of my photo editing involves humans, so basically this is just a useless product. I get that Google doesn't want to be responsible for deep fake advances, but that seems inevitable, so this is just slightly delay…

I've done ~20 prompts so far and not had one be rejected so far. What sort of things are you asking it to do? I've tried things like changing clothing and accessories on people.

Basic things like: "{uploaded image of a man} can you remove the glasses?" or "make everyone in the picture smile" or "open the eyes of everyone in the photo". Nothing that a human would consider "unsafe". I am based in EU and using Google AI Studio with all safety toggles set to "Off".

Re: Gemini 2.5 Flash Image

#434

I can imagine an automated blackmail bot that scrapes image, video, voice samples from anyone with the most meagre online presence, which then creates high resolution videos of that person doing the most horrid acts, then threatening to share those videos with that person's family, friends and business contacts unless they are paid $5000 in a cryptocurrency to an anonymous address. And further, I can imagine some per…

But these new amazing AI image generators lets you just say "It wasn't me, it is an AI fake". Long term they will seriously devalue blackmail material. I read a scifi novel where they invented a wormhole that only light could pass through but it could be used as a camera that could go anywhere and eventually anytime and there was absolutely no way to block it. So some people adapted to this fact by not wearing clothe…

The light of other days, by Arthur C. Clarke and Stephen Baxter. Really cool book.

Re: Gemini 2.5 Flash Image

#435
They should have called it emacs-banana, just to piss more people off.

And then call the next model vim-banana and start the editor-banana wars.

In all seriousness though, I'm seeing a worrying trend where google is hijacking well-known project names from other domains more and more now, with the "accidental"(?) side-effect of diluting their searchability and discoverability online, to the point I can no longer believe it is mere coincidence (whether malicious or not is another story altogether of course, but even if not, this is still a worrying trend).

I mean, what's next? Gopher 2.5 GIMP Video aka sublime-apple?

Re: Gemini 2.5 Flash Image

#436

Earlier quoted context omitted.

What tools did you use to make those videos from the PG image?

I used a bunch of models in conjunction: - Midjourney (background) - Qwen Image (restyle PG) - Gemini 2.5 Flash (editing in PG) - Gemini 2.5 Flash (adding YC logo) - Kling Pro (animation) I didn't spend too much time correcting mistakes. I used a desktop model aggregation and canvas tool that I wrote [1] to iterate and structure the work. I'll be open sourcing it soon. [1] https://getartcraft.com

What is PG?

Re: Gemini 2.5 Flash Image

#437

Earlier quoted context omitted.

I used a bunch of models in conjunction: - Midjourney (background) - Qwen Image (restyle PG) - Gemini 2.5 Flash (editing in PG) - Gemini 2.5 Flash (adding YC logo) - Kling Pro (animation) I didn't spend too much time correcting mistakes. I used a desktop model aggregation and canvas tool that I wrote [1] to iterate and structure the work. I'll be open sourcing it soon. [1] https://getartcraft.com

What is PG?

Paul Graham, Y Combinator founder.

Re: Gemini 2.5 Flash Image

#438

Earlier quoted context omitted.

I used a bunch of models in conjunction: - Midjourney (background) - Qwen Image (restyle PG) - Gemini 2.5 Flash (editing in PG) - Gemini 2.5 Flash (adding YC logo) - Kling Pro (animation) I didn't spend too much time correcting mistakes. I used a desktop model aggregation and canvas tool that I wrote [1] to iterate and structure the work. I'll be open sourcing it soon. [1] https://getartcraft.com

What is PG?

In this context, it's Paul Graham, the head Y Combinator guy whose cartoon likeness appears in the generated video: https://news.ycombinator.com/user?id=pg

Re: Gemini 2.5 Flash Image

#439

Unfortunately, it suffers from the same safetyism than other many releases. Half of the prompts get rejected. How can you have character consistency if the model is forbidden from editing any human. And most of my photo editing involves humans, so basically this is just a useless product. I get that Google doesn't want to be responsible for deep fake advances, but that seems inevitable, so this is just slightly delay…

I have an old photo of my girlfriend with her cousin when they were young, wearing Christmas dresses in front of the tree, not long before they were separated to other sides of the world for decades now. The photo is itself low quality on top of the photo itself being physically beat up. So far no model is willing to clean it up :/

Open source models like Flux Kontext or Qwen image edit wouldn't refuse, but you need to either have a sufficiently strong GPU or get one in the cloud (not difficult nor expensive with services like runpod), then set up your own processing pipeline (again, not too difficult if you use ComfyUI). Results won't be SOTA, but they shouldn't be too far off.

Re: Gemini 2.5 Flash Image

#440
Is this truly "native" to Gemini 2.5 Flash as they call it? Isn't this a dedicated and different text-to-image model hooked up to Gemini 2.5 Flash with function calling? or do they somehow merge the weights of the two models whilst also not incurring in side effects like degradation in performance?
Post reply on HN