Live data from Hacker News

Gemini 2.5 Flash Image

developers.googleblog.com

461–470 of 504 posts

Re: Gemini 2.5 Flash Image

#461
post #60

This is the gpt 4 moment for image editing models. Nano banana aka gemini 2.5 flash is insanely good. It made a 171 elo point jump in lmarena! Just search nano banana on Twitter to see the crazy results. An example. https://x.com/D_studioproject/status/1958019251178267111

> This is the gpt 4 moment for image editing models. No it's not. We've had rich editing capabilities since gpt-image-1, this is just faster and looks better than the (endearingly? called) "piss filter". Flux Kontext, SeedEdit, and Qwen Edit are all also image editing models that are robustly capable. Qwen Edit especially. Flux Kontext and Qwen are also possible to fine tune and run locally. Qwen (and its video gen s…

I'm totally with you. Dismayed by all these fanbois.

Re: Gemini 2.5 Flash Image

#462
post #52
post #23

Earlier quoted context omitted.

brown uniform, red armband with swastika was the usual SA look in the 1920s.

Mh. Apparently like this if we ask AI: https://postimg.cc/xX9K3kLP ...

i wonder if there was a single black or asian member of this particular military group, presumably not

Re: Gemini 2.5 Flash Image

#463

Earlier quoted context omitted.

Are their models that have vector space that includes ideas, not just words/media but not entirely corporeal aspects? So when generating a video of someone playing a keyboard the model would incorporate the idea of repeating groups of 8 tones, which is a fixed ideational aspect which might not be strongly represented in words adjacent to "piano". It seems like models need help with knowing what should be static, or h…

How would you encode those ideas?

I don't know, in part that's why I asked ... I wonder if there's a way to provide a loosely-defined space.

Perhaps it's a second word-vector space that allows context defined associations? Maybe it just needs tighter association of piano_keyboard with 8-step_repetition??

Re: Gemini 2.5 Flash Image

#464

Earlier quoted context omitted.

I don't mean to be rude, but this sounds like natural selection doing its work.

I'm pretty successful with an above average IQ. It was very convincing, along with three other college grads (one a medical doctor).

Shit happens bro.

Re: Gemini 2.5 Flash Image

#466

So it doesn't allow to do anything with photos containing kids, right? Isn't it too much of a filter for such a thing? ChatGPT thankfully created Ghibli versions of everything I gave it.

yes - any children images seem to be banned. Can't create a children's book with reference image.

Re: Gemini 2.5 Flash Image

#467
post #87

Earlier quoted context omitted.

Hope it works well for you! In my eyes, one specific example they show (“Prompt: Restore photo”) deeply AI-ifies the woman’s face. Sure it’ll improve over time of course.

Another question/concern for me: if I restore an old picture of my Gramma, will my Gramma (or a Gramma that looks strikingly similar) ever pop up on other people's "give me a random Gramma" prompts?

It might show her for prompts of “show me the world’s best grandma” :)

On free tier, I’d essentially believe that to be the default behavior. In reality they might simply use your feedback and your text prompts instead. Certainly know free Google/OpenAI LLM usage entails prompts being used for research.

Edit: decent chance it would NOT directly integrate grandma into its training, but would try hard to use an offline model for any privacy concerns

Re: Gemini 2.5 Flash Image

#468

Earlier quoted context omitted.

If you compare to the amount of effort required in Photoshop to achieve the same results, still a vast improvement

Vibe coding might not be real, but vibe graphics design certainly is. https://imgur.com/a/internet-DWzJ26B Anyone can make images and video now.

I think much like coding, the top of the game is all the old stuff and a bunch of new stuff that is impossible to master without some real math or at least outlier mathematical intuition.

The old top of the game is available to more people (though mid level people trying to level up now face a headwind in a further decoupling of easily read signals and true taste, making the old way of developing good taste harder).

This stuff makes people who were already "master rate" who are also nontrivially sophisticated machine learning hobbyists minimum and drives their peak and frontier out, drives break even collaboration overhead down.

It's always been possible to DIY code or graphic design, it's always been possible to tell the efforts of dabblers and pros apart, and unlike many commodities? There is rarely a "good enough". In software this is because compute is finite and getting more out of it pays huge, uneven returns, in graphic design its because extreme quality work is both aesthetically pleasing as well as a mark of quality (imperfect but a statement someone will commit resources).

And it's just hard to see it being different in any field. Lawyers? Opposing counsel has the best AI, your lawyer better have it too. Doctors? No amount of health is "enough" (in general).

I really think HN in particular but to some extent all CNBC-adjacent news (CEO OnlyFans stuff of all categories) completely misses the forest (the gap between intermediate and advanced just skyrocketed) for the trees (space-filling commodity knowledge work just plummeted in price).

But "commodity knowledge work" was always kind of an oxymoron, David Graeber called such work "bullshit jobs". You kinda need it to run a massive deficit in an over-the-hill neoliberal society, it's part of the " shift from production to consumption" shell game. But it's a very recent, very brief thing that's already looking more than wobbly. Outside of that? Apprentices, journeymen, masters is the model that built the world.

AI enables a new even more extreme form of mastery, blurs the line between journeyman and dabbler, and makes taking on apprentices a much longer-term investment (one of many reasons the PRC seems poised to enjoy a brief hegemony before demographics do in the Middle Kingdom for good, in China, all the GPUs run Opus, none run GPT-5 or LLaMA Behemoth).

The thing I really don't get is why CEOs are so excited about this and I really begin to suspect they haven't as a group thought it through (Zuckerberg maybe has, he's offering Tulloch a billion): the kind of CEO that manages a big pile of "bullshit jobs"?

AI can do most of their job today. Claude Opus 4.1? It sounds like if a mid-range CEO was exhaustively researched and gaff immune. Ditto career machine politicians. AI non practitioner prognosticators. That crowd.

But the top graphic communications people and CUDA kernel authors? Now they have to master ComfyUI or whatever and the color theory to get anything from it that stands out.

This is not a democratizing thing. And I cannot see it accruing to the Zuckerberg side of the labor/capital divvy up without a truly durable police state. Zuck offering my old chums nation state salaries is an extreme and likely transitory thing, but we know exactly how software professional economics work when it buckets as "sorcery" and "don't bother": that's 1950 to whenever we mark the start of the nepohacker Altman Era, call it 2015. In that world good hackers can do whatever they want, whenever they want, and the money guys grit their teeth. The non-sorcery bucket has paper mache hack-magnet hackathon projects in it at a fraction of the old price. So disruption, wow.

Whether that's good or bad is a value judgement I'll save for another blog post (thank you for attending my TED Talk).

Re: Gemini 2.5 Flash Image

#469

Earlier quoted context omitted.

[flagged]

Come on, don't be mean. Imagine saying this in person to someone who just told you they got scammed. "You're just extremely gullible" is just so mean...show some empathy.

Anyone who trusts Musk enough to send him $15k doesn't deserve a bit of empathy.

Re: Gemini 2.5 Flash Image

#470
post #459

Earlier quoted context omitted.

Yes, there was gemini-2.0-flash-preview-image-generation before, which could generate and edit also. But weaker than the new one.

Thanks, I'd not realised that, which means I have no idea if the things I've done outside of the API are this new one or not. That does feel classic google.

Yes, there's a conflict between wanting to just provide the good stuff by default under a unified Gemini brand where you don't have to worry about model names, it just works, versus building hype for a specific model and then being unclear about whether you're using that one or not. The nano-banana name is unique and fun, and got some recognition on social media already, they should just make a page with that heading and a chatbox. But again, that would focus on the new image editor thing only, and they probably want to lure people into their whole ecosystem, to switch to Gemini in general, from competitors like ChatGPT.
Post reply on HN