I...worked on the detailed Nano Banana prompt engineering analysis for months ( https://news.ycombinator.com/item?id=45917875 )...and...Google just...Google released a new version. Nano Banana Pro should work with my gemimg package ( https://github.com/minimaxir/gemimg ) without pushing a new version by passing: g = GemImg(model="gemini-3-pro-image-preview") I'll add the new output resolutions and other features ASAP…
>> - Put a strawberry in the left eye socket. >>- Put a blackberry in the right eye socket. >> All five of the edits are implemented correctly This is a GREAT example of the (not so) subtle mistakes AI will make in image generation, or code creation, or your future knee surgery. The model placed the specified items in the eye sockets based on the viewers left/right; when we talk relative in this scenario we usually (…
Nano Banana Pro
281–290 of 718 posts
Re: Nano Banana Pro
#282Earlier quoted context omitted.
> it's absolutely the definitive manual How do you know Simon? It's certainly a blog post, with content about prompting in it. If your goal is to make generative art that uses specific IP, I wouldn't use it.
Do you know of a better document specifically about prompting Nano Banana?
Re: Nano Banana Pro
#283Earlier quoted context omitted.
>> - Put a strawberry in the left eye socket. >>- Put a blackberry in the right eye socket. >> All five of the edits are implemented correctly This is a GREAT example of the (not so) subtle mistakes AI will make in image generation, or code creation, or your future knee surgery. The model placed the specified items in the eye sockets based on the viewers left/right; when we talk relative in this scenario we usually (…
I don't know if that's so much a mistake as it is ambiguity though? To me, using the viewer's perspective in this case seems totally reasonable. Does it still use the viewer's perspective if the prompt specifies "Put a strawberry in the _patient's left eye_"? If it does, then you're onto something. Otherwise I completely disagree with this.
Re: Nano Banana Pro
#284Earlier quoted context omitted.
It’s subtly incorrect. R/w permissions for example are described incorrectly on some nodes.
Then the question becomes, can it incorporate targeted feedback, or is it a oneshot-or-bust affair? My experience is that ChatGPT is very good at iterating on text (prose, code) but fairly bad at iterating on images. It struggles to integrate small changes, choosing instead to start over from scratch, with wildly different results. Thinking especially here of architectural stuff, where it does a great job laying out…
I've used Claude to generate fairly simple icons and launch images for an iOS game and I make sure to have it start with SVG files since those can be defined as code first. This way it's easier to iterate on specific elements of the image (certain shapes need to be moved to a different position, color needs to be changed, text needs an update, etc.).
FWIW not sure how Nano Banana Pro works though.
Re: Nano Banana Pro
#285Re: Nano Banana Pro
#286"mountain dew themed pokemon" is the first search prompt I always try with new image models and Nano Banna Pro just gave me a green pikachu.
Other models do a much better job of creating something new.
Re: Nano Banana Pro
#287Earlier quoted context omitted.
>> - Put a strawberry in the left eye socket. >>- Put a blackberry in the right eye socket. >> All five of the edits are implemented correctly This is a GREAT example of the (not so) subtle mistakes AI will make in image generation, or code creation, or your future knee surgery. The model placed the specified items in the eye sockets based on the viewers left/right; when we talk relative in this scenario we usually (…
>This is a GREAT example of the (not so) subtle mistakes AI will make in image generation, or code creation, or your future knee surgery. The mistake is in the prompting (not enough information). The AI did the best it could "What's the biggest known planet" "Jupiter" "NO I MEANT IN THE UNIVERSE!"
Re: Nano Banana Pro
#288Earlier quoted context omitted.
> Knives (under a certain size) are not regulated. Guns are regulated in most countries. Atomic bombs are definitely regulated I don’t think this is a good comparison: knives are easy to produce, guns a bit harder, atomic bombs definitely harder. You should find something that is as easy to produce as a knife, but regulated.
>You should find something that is as easy to produce as a knife, but regulated. The DEA and ATF have entered the chat
Re: Nano Banana Pro
#289Not all examples they gave were like this. The example they gave of the word "Typography" would have fooled me as human-made. The infographics stood out though. I would have immediately noticed that the String of Turtles infographic was AI generated because of the stylistic choices. Same for the guide on how to make chai. I would be "suspicious" of the example they gave of the weather forecast but wouldn't immediately flag at as AI generated.
Similar note, earlier I was able to tell if something was AI generated right off the bat by noticing that it had a "Deviant Art" quality to it. My immediate guess is that certain sources of training data are over-represented.
Re: Nano Banana Pro
#290Earlier quoted context omitted.
I don't know if that's so much a mistake as it is ambiguity though? To me, using the viewer's perspective in this case seems totally reasonable. Does it still use the viewer's perspective if the prompt specifies "Put a strawberry in the _patient's left eye_"? If it does, then you're onto something. Otherwise I completely disagree with this.
“The right socket” can only be implied one way when talking about a body just like you only have one right hand despite the fact that it is on my left when looking at you.
Same language, opposite meaning because of a particular noun + context.
I think the only thing obvious here is that there is no obvious solution other than adding lots of clarification to your prompt.