Live data from Hacker News

Nano Banana Pro

blog.google

401–410 of 718 posts

Re: Nano Banana Pro

#401

I don't understand the excitement around generating and/or watching AI-produced videos. To me it's probably the single most uninteresting and boring thing related to AI that I can think of. What is the appeal?

Sometimes, an animation is the best way to convey information.

Re: Nano Banana Pro

#402
post #317

Something I find weird about AI image generation models is that even though they no longer produce weird "artifacts" that give away that the fact that it was AI generated, you can still recognize that it's AI due to stylistic choices. Not all examples they gave were like this. The example they gave of the word "Typography" would have fooled me as human-made. The infographics stood out though. I would have immediately…

We are just very sharp when it comes to seeing small differences in images. I'm reminded of when the air force decided to create a pilot seat that worked for everyone. They took the average body dimensions of all their recruits and designed a seat to fit the average. It turned out, the seat fit none of their recruits. [1] I think AI image generation is a lot like this. When you train on all images, you get to this we…

What determines which “average” AI models latch onto? At a pixel level, the average of every image is a grayish rectangle; that's obviously not what we mean and AI does not produce that. At a slightly higher level, the average of every image is the average of every subject every photographed or drawn (human, tree, house, plate of food, ...) in concept space; but AI still doesn't generate a human with branches or a house with spaghetti on it. At a still higher level there are things we recognize as sensible scenes, e.g., barista pouring a cup of coffee, anime scene of a guy fighting a robot, watercolor of a boat on a lake, which AI still does not (by default) average into, say, an equal parts watercolor/anime/photorealistic image of a barista fighting a robot on a boat while pouring a cup of coffee.

But it is undeniable that AI images do have an “average” feel to them. What causes this? What is the space over which AI is taking an average to produce its output? One possible answer is that a finite model size means that the model can only explore image space with a limited resolution, and as models get bigger/better they can average over a smaller and smaller portion of this space, but it is always limited.

But that raises the question of why models don't just naturally land on a point in image space. Is this just a limitation of training, which punishes big failures more strongly than it rewards perfection? Or is there something else at play here that's preventing models from landing directly on a “real” image?

Re: Nano Banana Pro

#403

Google has been stomping around like Godzilla this week, and this is the first time I decided to link my card to their AI studio. I had seen people saying that they gave up and went to another platform because it was "impossible to pay". I thought this was strange, but after trying to get a working API key for the past half hour, I see what they mean. Everything is set up, I see a message that says "You're using Paid…

100% this. I am using the pro/max plans on both claude and openai. Would love to experiment with gemini but paying is next to impossible. Why do i need the risk of a full blown gcp project just to test gemini. No thx.

[deleted]

Re: Nano Banana Pro

#404

Earlier quoted context omitted.

>> - Put a strawberry in the left eye socket. >>- Put a blackberry in the right eye socket. >> All five of the edits are implemented correctly This is a GREAT example of the (not so) subtle mistakes AI will make in image generation, or code creation, or your future knee surgery. The model placed the specified items in the eye sockets based on the viewers left/right; when we talk relative in this scenario we usually (…

>This is a GREAT example of the (not so) subtle mistakes AI will make in image generation, or code creation, or your future knee surgery. The mistake is in the prompting (not enough information). The AI did the best it could "What's the biggest known planet" "Jupiter" "NO I MEANT IN THE UNIVERSE!"

asking "x-most known y" and not expecting a global answer is odd

Re: Nano Banana Pro

#405

Earlier quoted context omitted.

It doesn't affect your point but technically since the IAU are insane, exoplanets aren't technically planets and Jupiter is the largest planet in the universe.

I suppose it was too much to hope that chatbots could be trained to avoid pointless pedantry.

They've been trained on every web forum on the Internet. How could it be possible for them to avoid that?

Re: Nano Banana Pro

#406

Google has been stomping around like Godzilla this week, and this is the first time I decided to link my card to their AI studio. I had seen people saying that they gave up and went to another platform because it was "impossible to pay". I thought this was strange, but after trying to get a working API key for the past half hour, I see what they mean. Everything is set up, I see a message that says "You're using Paid…

I had to write a post request to try it when it launched

Re: Nano Banana Pro

#407

Alright results are in! I've re-run all my editing based adherence related prompts through Nano Banana Pro. NB Pro managed to successfully pass SHRDLU, the M&M Van Halen test (as verified independently by Simon), and the Scorpio street test - all of which the original NB failed. Model results 1. Nano Banana Pro: 10 / 12 2. Seedream4: 9 / 12 3. Nano Banana: 7 / 12 4. Qwen Image Edit: 6 / 12 https://genai-showdown.spec…

thanks, I love your website. Are you planning to do NB Pro for the text-to-image benchmark too?

Re: Nano Banana Pro

#408

Google has been stomping around like Godzilla this week, and this is the first time I decided to link my card to their AI studio. I had seen people saying that they gave up and went to another platform because it was "impossible to pay". I thought this was strange, but after trying to get a working API key for the past half hour, I see what they mean. Everything is set up, I see a message that says "You're using Paid…

You can use it also in Gemini.

I hate that they kinda try to hide the model version. Like if you click the dropdown in the chat box, you can see that "Thinking" means 3 Pro. When you select the "Create images" tool, it doesn't tell you it's using Nano Banana Pro until it actually starts generating the image.

Tell me the model it's using. It's as if Google is trying to unburden me with the knowledge of what model does what but it's just making things more confusing.

Oh, and setting up AI Studio is a mess. First I have to create a project. Then an API key. Then I have to link the API key to the project. Then I have to link the project to the chat session... Come on, Google.

Re: Nano Banana Pro

#409
post #364

Earlier quoted context omitted.

How do you disagree with having a right and a left hand?

GP is using right as in “correct”, not directionality.

No, I don't think they are.

If you are facing a wall-plate with two power sockets on it side by side and you are telling someone to plug something in, which one would be "the right socket", and which would be "the left socket"?

If above the wall-plate is a photo of a person and you are someone to draw a tattoo on the photo, which is "the right arm" and which is "the left arm"?

Same wording, different expectation.

Re: Nano Banana Pro

#410

Alright results are in! I've re-run all my editing based adherence related prompts through Nano Banana Pro. NB Pro managed to successfully pass SHRDLU, the M&M Van Halen test (as verified independently by Simon), and the Scorpio street test - all of which the original NB failed. Model results 1. Nano Banana Pro: 10 / 12 2. Seedream4: 9 / 12 3. Nano Banana: 7 / 12 4. Qwen Image Edit: 6 / 12 https://genai-showdown.spec…

thanks, I love your website. Are you planning to do NB Pro for the text-to-image benchmark too?

Definitely! Even though NB's predominant use case seems to be editing, it's still producing surprisingly decent text-to-image results. Imagen4 currently still comes out ahead in terms of image fidelity, but I think NB Pro will close the gap even further.

I'll try to have the generative comparisons for NB Pro up later this afternoon once I catch my breath.

Post reply on HN