Live data from Hacker News

Gemini 2.5 Flash Image

developers.googleblog.com

231–240 of 504 posts

Re: Gemini 2.5 Flash Image

#231
"Can you make a version of this picture where I wear the best possible sunglasses for my face shape?"

made me realize that AI image modification is now technically flawless, utterly devoid of taste, and that I myself am a rather unattractive fellow.

Re: Gemini 2.5 Flash Image

#232
post #134

Earlier quoted context omitted.

hm, but isn't it wild thinking that elon is talking to you and asking you for 15k , like bro has the money of his lifetime, why would he ask you? It doesn't make that much sense idk

I remember watching the SpaceX channel on youtube, which isn't a legit source. AI Elon basically says "I want to help make bitcoin more popular, let me show you how easy it it to transfer money around with btc. Send my $X and I'll send you back $2X! It's very inline with a typical elon message (I'll give you 1 million to vote R), it's on a channel called SpaceX. It's pretty believable. Granted I played Runescape and…

It's only believable to the extent that I believe that Musk would actually run such a transparently obvious scam.

Re: Gemini 2.5 Flash Image

#233

I digitised our family photos but a lot of them were damaged (shifted colours, spills, fingerprints on film, spots) that are difficult to correct for so many images. I've been waiting for image gen to catch up enough to be able to repair them all in bulk without changing details, especially faces. This looks very good at restoring images without altering details or adding them where they are missing, so it might fina…

I don't really understand the point of this usecase. Like, can't you also imagine what the photos might look like without the damage? Same with AI upscaling in phone cameras... if I want a hypothetical idea of what something in the distance might look like, I can just... imagine it?

I think we will eventually have AI based tools that are just doing what a skilled human user would do in Photoshop, via tool-use. This would make sense to me. But just having AI generate a new image with imagined details just seems like waste of time.

Re: Gemini 2.5 Flash Image

#234

Like most image generators, it didn’t pass the piano keyboard test. (Black keys are wrong.) https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%...

Like most image models, except GPT-4o, it also didn't pass the wooden Penrose triangle test. (It creates normal triangles.)

Re: Gemini 2.5 Flash Image

#235

I am glad that I never decided to become a photoshop pro. I always contemplated about it, seemed attractive for a while, but glad that I decided against it. RIP r/photoshopbattles. It was in the endless list of new shiny 'skills' that feels good to have. Now I can use nano-banana instead. Other models will soon follow, I am sure.

[dead]

Re: Gemini 2.5 Flash Image

#236
post #217

Earlier quoted context omitted.

[flagged]

I got scammed similarly (although $10, because I tested first), because 1. it was on YouTube, on a channel called "SpaceX" with verified logo 2. with hundreds of thousands of viewers live 3. with a believable speech from Mr. Musk standing next to its rockets (and knowing his interest in cryptocurrencies). This happened as I was genuinely searching for the actual live stream of SpaceX. I am ashamed, even more so becau…

I am flabbergasted that you both get scammed. I would understand if this was two years ago, but now? Do people really not know about these scams? I can already see down votes coming for victim blaming, but this is to me really shocking. Notice that there isn't "tell hn: don't get scammed by deep fake crypto Elon" because people who usually posts also consider this general knowledge. That's why it's so effective I guess. In a similar manner there will never be "tell hn: don't drink acid it will burn your intestines", the danger is so obvious that nobody feels the need to post it and because nobody is posting it, people get scammed. I don't know what is the solution to that. How should you tell people what everybody should be already knowing?

I remember being on a machining workshop and he was telling such an obvious things. Obvious things are obvious until they aren't, and then somebody gets hurt.

Re: Gemini 2.5 Flash Image

#237

Like most image generators, it didn’t pass the piano keyboard test. (Black keys are wrong.) https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%...

Are their models that have vector space that includes ideas, not just words/media but not entirely corporeal aspects?

So when generating a video of someone playing a keyboard the model would incorporate the idea of repeating groups of 8 tones, which is a fixed ideational aspect which might not be strongly represented in words adjacent to "piano".

It seems like models need help with knowing what should be static, or homomorphic, across or within images associated with the same word vectors and that words alone don't provide a strong enough basis [*1] for this.

*1 - it's so hard to find non-conflicting words, obviously I don't mean basis as in basis vectors, though there is some weak analogy.

Re: Gemini 2.5 Flash Image

#238
post #60

This is the gpt 4 moment for image editing models. Nano banana aka gemini 2.5 flash is insanely good. It made a 171 elo point jump in lmarena! Just search nano banana on Twitter to see the crazy results. An example. https://x.com/D_studioproject/status/1958019251178267111

Completely agree - I make logos for my github projects for fun, and the last time I tried SOTA image generation for logos, it was consistently ignoring instructions and not doing anything close to what i was asking for. Google's new release today did it near flawlessly, exactly how I wanted it, in a single prompt. A couple more prompts for tweaking (centering it, rotating it slightly) got it perfect. This is awesome.

Re: Gemini 2.5 Flash Image

#239

Like most image generators, it didn’t pass the piano keyboard test. (Black keys are wrong.) https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%...

or my "hands with palms facing down" test.... no matter how hard I try it just can't get open hands, palms down.

I guess the vast majority of images have the palms the other way, that this biases the output. It's like how we misinterpret images to generate optical illusions, because we're expecting valid 3D structures (Escher's staircases, say).
Post reply on HN