Live data from Hacker News

Nano Banana Pro

blog.google

291–300 of 718 posts

Re: Nano Banana Pro

#291
post #266

This is the first image model I’ve used that passed my piano test. It actually generated an image of a keyboard with the proper pattern of black keys repeated per octave – every other model I’ve tried this with since the first Dall-E has struggled to render more than a single octave, usually clumping groups of two black keys or grouping them four at a time. Very impressive grasp of recursive patterns.

If you ask it for anything outside of the standard 88 key set it falls short. For instance "Generate a piano, but have the left most key start at middle C, and the notes continue in the standard order up (D, E, F, G, ...) to the right most key" The above prompt will be wrong, seemingly every time. The model has no understanding of the keys or where they belong, and it is not able to intuit creating something within t…

Yep - one of my goto bench marks is a "historical piano" - meaning the naturals are black and the sharps/flats are white.

https://imgur.com/a/SZbzsYv

Re: Nano Banana Pro

#292

Earlier quoted context omitted.

It’s subtly incorrect. R/w permissions for example are described incorrectly on some nodes.

Then the question becomes, can it incorporate targeted feedback, or is it a oneshot-or-bust affair? My experience is that ChatGPT is very good at iterating on text (prose, code) but fairly bad at iterating on images. It struggles to integrate small changes, choosing instead to start over from scratch, with wildly different results. Thinking especially here of architectural stuff, where it does a great job laying out…

Nano Banana is really good at iterating on images, as shown by the pancake skull example I borrowed from Max Woolf: https://simonwillison.net/2025/Nov/20/nano-banana-pro/#tryin...

I've tried iterating on slides with test on them a bit and it seems to be competent at that too.

Re: Nano Banana Pro

#293

Earlier quoted context omitted.

>This is a GREAT example of the (not so) subtle mistakes AI will make in image generation, or code creation, or your future knee surgery. The mistake is in the prompting (not enough information). The AI did the best it could "What's the biggest known planet" "Jupiter" "NO I MEANT IN THE UNIVERSE!"

No, this is squarely on the AI. A human would know what you mean without specific instructions.

I would not, I would clarify, and I think I'm a human.

Re: Nano Banana Pro

#294

Earlier quoted context omitted.

“The right socket” can only be implied one way when talking about a body just like you only have one right hand despite the fact that it is on my left when looking at you.

"Plug into right power socket" Same language, opposite meaning because of a particular noun + context. I think the only thing obvious here is that there is no obvious solution other than adding lots of clarification to your prompt.

I think you missed the entire point?

Re: Nano Banana Pro

#295
post #213

I...worked on the detailed Nano Banana prompt engineering analysis for months ( https://news.ycombinator.com/item?id=45917875 )...and...Google just...Google released a new version. Nano Banana Pro should work with my gemimg package ( https://github.com/minimaxir/gemimg ) without pushing a new version by passing: g = GemImg(model="gemini-3-pro-image-preview") I'll add the new output resolutions and other features ASAP…

In case anyone missed Max's Nano Banana prompting guide, it's absolutely the definitive manual for prompting the original Nano Banana... and I tried some of the prompts in there against Nano Banana Pro and found it to be very applicable to the new model as well. https://minimaxir.com/2025/11/nano-banana-prompts/#hello-nan... My recreations of those pancake batter skulls using Nano Banana Pro: https://simonwillison.ne…

In my experience multimodal models like gpt-image-1/nano/etc. don't really require a lot of prompt trickery [1] like the good ol' days of SD 1.5.

To be clear, that's a good thing though. It's also one of the reasons why "prompt engineering" will become less relevant as model understanding goes up.

[1] - Unless you're trying to circumvent guardrails

Re: Nano Banana Pro

#296
post #270

Earlier quoted context omitted.

Do you know of a better document specifically about prompting Nano Banana?

Why don't you just ask Gemini? It will tell you! There's no mystery.

You implied that Max's Nano Banana prompting guide wasn't the best available, so I think it's on you to provide a link to a better one.

Re: Nano Banana Pro

#297
post #162

Earlier quoted context omitted.

And that's not a good thing.

Why not? Like, genuinely.

I generally don't think that's it's good or just for a government to collude with manufacturers to track/trace it's citizens without consent or notice. And even if notice was given, I'd still be against it

The arguments put forward by people generally I don't find compelling -- for example, in this thread around protecting against counterfeit.

The "force" applied to address these concerns is totally out of proportion. Whenever these discussions happen, I feel like they descend into a general viewpoint, "if we could technically solve any possible crime, we should do everything in our power to solve it."

I'm against this viewpoint, and acknowledge that that means _some crime_ occurs. That's acceptable to me. I don't feel that society is correctly structured to "treat" crime appropriately, and technology has outpaced our ability to holistically address it.

Generally, I don't see (speaking for the US) the highest incarceration rate in the world to be a good thing, or being generally effective, and I don't believe that increasing that number will change outcomes.

Re: Nano Banana Pro

#298

You can try it out for free on LMArena [0]: New Chat -> Battle dropdown -> Direct Chat -> Click on Generate Image in the chat box -> Click dropdown from hunyuan-image-3.0 -> gemini-3-pro-image-preview (nano-banana-pro). I've only managed to get a few prompts to go through, if it takes longer than 30 seconds it seems to just time out. Image quality seems to vary wildly; the first image I tried looked really good but t…

Thanks - this worked for me (some errors, some success).

Last week I was making a birthday card for my son with the old model. The new model is dramatically better - I'm asking for an image in comic book style, prompted with some images of him.

With the previous model, the boy was descriptively similar (e.g. hair colour and style) but looked nothing like him. With this model it's recognisably him.

Re: Nano Banana Pro

#299

Something I find weird about AI image generation models is that even though they no longer produce weird "artifacts" that give away that the fact that it was AI generated, you can still recognize that it's AI due to stylistic choices. Not all examples they gave were like this. The example they gave of the word "Typography" would have fooled me as human-made. The infographics stood out though. I would have immediately…

I think it's because they're all trained on the same data (everything they could possibly scrape from the open web). The models tend to learn some kind of distribution of what is most likely for a given prompt. It tends to produce things that are very average looking, very "likely", but as a result also predictable and unoriginal.

If you want something that looks original, you have to come up with a more original prompt. Or we have to find a way to train these models to sample things that are less likely from their distribution? Find a way to mathematically describe what it means to be original.

Re: Nano Banana Pro

#300
post #94

Earlier quoted context omitted.

I wouldn’t call LLMs a dead end, they’re so useful as-is

LLMs are useful, but they've hit a wall on the path to automating our jobs. Benchmark scores are just getting better at test taking. I don't see them replacing software engineers without overcoming obstacles. AI for images, video, music - these tools can already make movies, games, and music today with just a little bit of effort by domain experts. They're 10,000x time and cost savers. The models and tools are contin…

I'm literally a software engineer, and a business owner. I don't think about this in binary terms (replacement or not), but just like CMS's replaced the jobs of people that write HTML by hand to build websites, I think whole classes of software development will get democratized.

For example, I'm currently vibe coding an app that will be specific to our company, that helps me run all the aspects of our business and integrates with our systems (so it'll integrate with quickbooks for invoicing, etc), and help us track whether we have the right insurance across multiple contracts, will remind me about contract deadlines coming up, etc.

It's going to combine the information that's currently in about 10 different slightly out of sync spreadsheets, about 2 dozen google docs/drive files, and multiple external systems (Gusto, Quickbooks, email, etc).

Even though I could build all this manually (as a software developer), I'd never take the time to do it, because it takes away from client work. But now I can actually do it because the pace is 100x faster, and in the background while I'm doing client work.

Post reply on HN