Live data from Hacker News

Nano Banana Pro

blog.google

21–30 of 718 posts

Re: Nano Banana Pro

#21

> Generate better visuals with more accurate, legible text directly in the image in multiple languages Assuming that this new model works as advertised, it's interesting to me that it took this long to get an image generation model that can reliably generate text. Why is text generation in images so hard?

It’s not necessarily harder than other aspects. However:

- It requires an AI that actually understands English, I.e. an LLM. Older, diffusion-only models were naturally terrible at that, because they weren’t trained on it.

- It requires the AI to make no mistakes on image rendering, and that’s a high bar. Mistakes in image generation are so common we have memes about it, and for all that hands generally work fine now, the rest of the picture is full of mistakes you can’t tell are mistakes. Entirely impossible with text.

Nano Banana Pro seems to somewhat reliably produce entire pictures without any mistakes at all.

Re: Nano Banana Pro

#22
post #3

Can anyone please explain me the invisible watermarking mentioned in the said promo?

It's called Synth ID. It's a watermark that proves an image was generated by AI. https://deepmind.google/models/synthid/

So whoever creates AI content needs to voluntarily adopt this so that Google can sell "technology" for identifying said content?

Not sure how that makes any sense

Re: Nano Banana Pro

#25

Earlier quoted context omitted.

It's called Synth ID. It's a watermark that proves an image was generated by AI. https://deepmind.google/models/synthid/

Super important for Google as a search engine so they can filter out and downrank AI generated results. However I expect there are many models out there which don’t do this, that everyone could use instead. So in the end a “feature” like this makes me less likely to use their model because I don’t know how Google will end up treating my blog post if I decide to include an AI generated or AI edited image.

[deleted]

Re: Nano Banana Pro

#27

I guess the true endgame of AI products is naming them. We still have quite a way to go.

This has always been the hardest problem in computer science besides “Assume a lightweight J2EE distribution…”

Re: Nano Banana Pro

#28

> Generate better visuals with more accurate, legible text directly in the image in multiple languages Assuming that this new model works as advertised, it's interesting to me that it took this long to get an image generation model that can reliably generate text. Why is text generation in images so hard?

As a complete layman, it seems obvious that it should be hard? Like, text is a type of graphic that needs to be coherent both in its detail and its large structure, and there’s a very small amount of variation that we don’t immediately notice as strange or flat out incorrect. That’s not true of most types of imagery.

Re: Nano Banana Pro

#30

Earlier quoted context omitted.

It's called Synth ID. It's a watermark that proves an image was generated by AI. https://deepmind.google/models/synthid/

Super important for Google as a search engine so they can filter out and downrank AI generated results. However I expect there are many models out there which don’t do this, that everyone could use instead. So in the end a “feature” like this makes me less likely to use their model because I don’t know how Google will end up treating my blog post if I decide to include an AI generated or AI edited image.

It’s required by EU regulations. Any public generator that doesn’t do it, is in violation of that unless it’s entirely inaccessible from the EU…

But of course there’s no way to enforce it on local generation.

Post reply on HN