Live data from Hacker News

Nano Banana Pro

blog.google

281–290 of 718 posts

Re: Nano Banana Pro

#281

I...worked on the detailed Nano Banana prompt engineering analysis for months ( https://news.ycombinator.com/item?id=45917875 )...and...Google just...Google released a new version. Nano Banana Pro should work with my gemimg package ( https://github.com/minimaxir/gemimg ) without pushing a new version by passing: g = GemImg(model="gemini-3-pro-image-preview") I'll add the new output resolutions and other features ASAP…

>> - Put a strawberry in the left eye socket. >>- Put a blackberry in the right eye socket. >> All five of the edits are implemented correctly This is a GREAT example of the (not so) subtle mistakes AI will make in image generation, or code creation, or your future knee surgery. The model placed the specified items in the eye sockets based on the viewers left/right; when we talk relative in this scenario we usually (…

Right, that's why one should use "put a strawberry in the portside eye socket" and "put a strawberry in the starboard side socket"

Re: Nano Banana Pro

#282
post #270

Earlier quoted context omitted.

> it's absolutely the definitive manual How do you know Simon? It's certainly a blog post, with content about prompting in it. If your goal is to make generative art that uses specific IP, I wouldn't use it.

Do you know of a better document specifically about prompting Nano Banana?

Why don't you just ask Gemini? It will tell you! There's no mystery.

Re: Nano Banana Pro

#283
post #275

Earlier quoted context omitted.

>> - Put a strawberry in the left eye socket. >>- Put a blackberry in the right eye socket. >> All five of the edits are implemented correctly This is a GREAT example of the (not so) subtle mistakes AI will make in image generation, or code creation, or your future knee surgery. The model placed the specified items in the eye sockets based on the viewers left/right; when we talk relative in this scenario we usually (…

I don't know if that's so much a mistake as it is ambiguity though? To me, using the viewer's perspective in this case seems totally reasonable. Does it still use the viewer's perspective if the prompt specifies "Put a strawberry in the _patient's left eye_"? If it does, then you're onto something. Otherwise I completely disagree with this.

“The right socket” can only be implied one way when talking about a body just like you only have one right hand despite the fact that it is on my left when looking at you.

Re: Nano Banana Pro

#284

Earlier quoted context omitted.

It’s subtly incorrect. R/w permissions for example are described incorrectly on some nodes.

Then the question becomes, can it incorporate targeted feedback, or is it a oneshot-or-bust affair? My experience is that ChatGPT is very good at iterating on text (prose, code) but fairly bad at iterating on images. It struggles to integrate small changes, choosing instead to start over from scratch, with wildly different results. Thinking especially here of architectural stuff, where it does a great job laying out…

I would assume it depends on how it generates the images.

I've used Claude to generate fairly simple icons and launch images for an iOS game and I make sure to have it start with SVG files since those can be defined as code first. This way it's easier to iterate on specific elements of the image (certain shapes need to be moved to a different position, color needs to be changed, text needs an update, etc.).

FWIW not sure how Nano Banana Pro works though.

Re: Nano Banana Pro

#285
Everyone who worked on this is a traitor to the human race. Why do we need to make it impossible to make a living as an artist? Who thinks an endless tsunami of garbage “content” churned out by machines dropping the bottom out of all artistic disciplines is a good idea?

Re: Nano Banana Pro

#286
There's some really impressive things about this (the speed, the lack of typical AI image gen artifacts) but it also seems less creative than other models I've tried?

"mountain dew themed pokemon" is the first search prompt I always try with new image models and Nano Banna Pro just gave me a green pikachu.

Other models do a much better job of creating something new.

Re: Nano Banana Pro

#287

Earlier quoted context omitted.

>> - Put a strawberry in the left eye socket. >>- Put a blackberry in the right eye socket. >> All five of the edits are implemented correctly This is a GREAT example of the (not so) subtle mistakes AI will make in image generation, or code creation, or your future knee surgery. The model placed the specified items in the eye sockets based on the viewers left/right; when we talk relative in this scenario we usually (…

>This is a GREAT example of the (not so) subtle mistakes AI will make in image generation, or code creation, or your future knee surgery. The mistake is in the prompting (not enough information). The AI did the best it could "What's the biggest known planet" "Jupiter" "NO I MEANT IN THE UNIVERSE!"

No, this is squarely on the AI. A human would know what you mean without specific instructions.

Re: Nano Banana Pro

#288
post #158

Earlier quoted context omitted.

> Knives (under a certain size) are not regulated. Guns are regulated in most countries. Atomic bombs are definitely regulated I don’t think this is a good comparison: knives are easy to produce, guns a bit harder, atomic bombs definitely harder. You should find something that is as easy to produce as a knife, but regulated.

>You should find something that is as easy to produce as a knife, but regulated. The DEA and ATF have entered the chat

They can leave, plain water fits this bill.

Re: Nano Banana Pro

#289
Something I find weird about AI image generation models is that even though they no longer produce weird "artifacts" that give away that the fact that it was AI generated, you can still recognize that it's AI due to stylistic choices.

Not all examples they gave were like this. The example they gave of the word "Typography" would have fooled me as human-made. The infographics stood out though. I would have immediately noticed that the String of Turtles infographic was AI generated because of the stylistic choices. Same for the guide on how to make chai. I would be "suspicious" of the example they gave of the weather forecast but wouldn't immediately flag at as AI generated.

Similar note, earlier I was able to tell if something was AI generated right off the bat by noticing that it had a "Deviant Art" quality to it. My immediate guess is that certain sources of training data are over-represented.

Re: Nano Banana Pro

#290
post #275

Earlier quoted context omitted.

I don't know if that's so much a mistake as it is ambiguity though? To me, using the viewer's perspective in this case seems totally reasonable. Does it still use the viewer's perspective if the prompt specifies "Put a strawberry in the _patient's left eye_"? If it does, then you're onto something. Otherwise I completely disagree with this.

“The right socket” can only be implied one way when talking about a body just like you only have one right hand despite the fact that it is on my left when looking at you.

"Plug into right power socket"

Same language, opposite meaning because of a particular noun + context.

I think the only thing obvious here is that there is no obvious solution other than adding lots of clarification to your prompt.

Post reply on HN