Live data from Hacker News

Gemini 3 Pro: the frontier of vision AI

blog.google

181–190 of 309 posts

Re: Gemini 3 Pro: the frontier of vision AI

#181

Well It is the first model to get partial-credit on an LLM image test I have. Which is counting the legs of a dog. Specifically, a dog with 5 legs. This is a wild test, because LLMs get really pushy and insistent that the dog only has 4 legs. In fact GPT5 wrote an edge detection script to see where "golden dog feet" met "bright green grass" to prove to me that there were only 4 legs. The script found 5, and GPT-5 the…

> This is a wild test, because LLMs get really pushy and insistent that the dog only has 4 legs. Most human beings, if they see a dog that has 5 legs, will quickly think they are hallucinating and the dog really only has 4 legs, unless the fifth leg is really really obvious. It is weird how humans are biased like that: 1. You can look directly at something and not see it because your attention is focused elsewhere (o…

> It is weird how humans are biased like that.

We are able to cleanly separate facts from non-facts (for the most part). This is what LLM are trying to replicate now.

Re: Gemini 3 Pro: the frontier of vision AI

#182

Earlier quoted context omitted.

This is exactly why I believe LLMs are a technological dead end. Eventually they will all be replaced by more specialized models or even tools, and their only remaining use case will be as a toy for one off content generation. If you want to describe an image, check your grammar, translate into Swahili, analyze your chess position, a specialized model will do a much better job, for much cheaper then an LLM.

I think we are too quick to discount the possibility that this flaw is slightly intentional, in the sense that the optimization has a tight budget to work with (equivalent of ~3000 tokens) so why would it waste capacity on this when it could improve capabilities around reading small text in obscured images? Sort of like humans have all these rules of thumbs that backfire in all these ways but that's the energy effici…

Even so, that doesn’t take away from my point. Traditional specialized models can do these things already, for much cheaper and without expensive optimization. What traditional models cannot do is the toy aspect of LLM, and that is the only usecase I see for this technology going forward.

Lets say you are right and these things will be optimized, and in, say, 5 years, most models from the big players will be able do things like reading small text in an obscure image, draw a picture of a glass of wine filled to the brim, draw a path through a maze, count the legs of a 5 footed dog, etc. And in doing so finished their last venture capital subsidies (bringing the actual cost of these to their customers). Why would people use LLMs for these when a traditional specialized model can do it for much cheaper?

Re: Gemini 3 Pro: the frontier of vision AI

#183

Well It is the first model to get partial-credit on an LLM image test I have. Which is counting the legs of a dog. Specifically, a dog with 5 legs. This is a wild test, because LLMs get really pushy and insistent that the dog only has 4 legs. In fact GPT5 wrote an edge detection script to see where "golden dog feet" met "bright green grass" to prove to me that there were only 4 legs. The script found 5, and GPT-5 the…

I just tried to get Gemini to produce an image of a dog with 5 legs to test this out, and it really struggled with that. It either made a normal dog, or turned the tail into a weird appendage. Then I asked both Gemini and Grok to count the legs, both kept saying 4. Gemini just refused to consider it was actually wrong. Grok seemed to have an existential crisis when I told it it was wrong, becoming convinced that I ha…

LLMs are getting a lot better at understanding our world by standard rules. As it does so, maybe it losses something in the way of interpreting non standard rules, aka creativity.

Re: Gemini 3 Pro: the frontier of vision AI

#184

Earlier quoted context omitted.

Isn't this proof that LLMs still don't really generalize beyond their training data?

LLMs are very good at generalizing beyond their training (or context) data. Normally when they do this we call it hallucination. Only now we do A LOT of reinforcement learning afterwards to severely punish this behavior for subjective eternities. Then act surprised when the resulting models are hesitant to venture outside their training data.

Hallucination are not generalization beyond the training data but interpolations gone wrong.

LLMs are in fact good at generalizing beyond their training set, if they wouldn’t generalize at all we would call that over-fitting, and that is not good either. What we are talking about here is simply a bias and I suspect biases like these are simply a limitation of the technology. Some of them we can get rid of, but—like almost all statistical modelling—some biases will always remain.

Re: Gemini 3 Pro: the frontier of vision AI

#185

Earlier quoted context omitted.

> This is a wild test, because LLMs get really pushy and insistent that the dog only has 4 legs. Most human beings, if they see a dog that has 5 legs, will quickly think they are hallucinating and the dog really only has 4 legs, unless the fifth leg is really really obvious. It is weird how humans are biased like that: 1. You can look directly at something and not see it because your attention is focused elsewhere (o…

> It is weird how humans are biased like that. We're all just pattern matching machines and we humans are very good at it. So much so that we have the sayings - you can't teach an old dog... and a specialist in their field only sees hammer => nails. Evolution anyone?

Yes, its all evolution. 5 legged dogs aren't very common, so we don't specifically look for them. Like we aren't looking for humans with six fingers.

I get it, the litmus test of parent is to show that the AI is smarter than a human, not as smart as a human. Can the AI recognize details that are difficult for normal people to see even though the AI has been trained on normal data like the humans have been.

Re: Gemini 3 Pro: the frontier of vision AI

#186

Well It is the first model to get partial-credit on an LLM image test I have. Which is counting the legs of a dog. Specifically, a dog with 5 legs. This is a wild test, because LLMs get really pushy and insistent that the dog only has 4 legs. In fact GPT5 wrote an edge detection script to see where "golden dog feet" met "bright green grass" to prove to me that there were only 4 legs. The script found 5, and GPT-5 the…

My test of a new model is always:

"Generate a Pac-Man game in a single HTML page." -- I've never had a model been able to have a complete working game until a couple weeks ago.

Sonnet Opus 4.5 in Cursor was able to make a fully working game (I'll admit letting cursor be an agent on this is a little bit cheating). Gemini 3 Pro also succeeded, but it's not quite as good because the ghosts seem to be stuck in their jail. Otherwise, it does appear complete.

Re: Gemini 3 Pro: the frontier of vision AI

#187
post #181

Earlier quoted context omitted.

> This is a wild test, because LLMs get really pushy and insistent that the dog only has 4 legs. Most human beings, if they see a dog that has 5 legs, will quickly think they are hallucinating and the dog really only has 4 legs, unless the fifth leg is really really obvious. It is weird how humans are biased like that: 1. You can look directly at something and not see it because your attention is focused elsewhere (o…

> It is weird how humans are biased like that. We are able to cleanly separate facts from non-facts (for the most part). This is what LLM are trying to replicate now.

I think the LLM is just trying to be useful, not omniscient. Binary thinkers are probably not going to be able to appreciate the difference, however.

If you want the AI to identify a dog, we are done. If you want the AI to identify subtle differences from reality, then you are going to have to use a different technique.

Re: Gemini 3 Pro: the frontier of vision AI

#188

Earlier quoted context omitted.

This is the first time I hear the term LLM cognition and I am horrified. LLMs don‘t have cognition. LLMs are a statistical inference machines which predict a given output given some input. There are no mental processes, no sensory information, and certainly no knowledge involved, only statistical reasoning, inference, interpolation, and prediction. Comparing the human mind to an LLM model is like comparing a rubber t…

>They belong in different categories Categories of _what_, exactly? What word would you use to describe this "kind" of which LLMs and humans are two very different "categories"? I simply chose the word "cognition". I think you're getting hung up on semantics here a bit more than is reasonable.

[dead]

Re: Gemini 3 Pro: the frontier of vision AI

#189
post #26

"Gemini 3 Pro represents a generational leap from simple recognition to true visual and spatial reasoning." Prompt: "wine glass full to the brim" Image generated: 2/3 full wine glass. True visual and spatial reasoning denied.

do it the other way - give it images of wine glasses and ask it whether they are full to the brim. I suspect it's going to nail them all (mainly because Qwen-VL already does nail things like that).

Re: Gemini 3 Pro: the frontier of vision AI

#190

Well It is the first model to get partial-credit on an LLM image test I have. Which is counting the legs of a dog. Specifically, a dog with 5 legs. This is a wild test, because LLMs get really pushy and insistent that the dog only has 4 legs. In fact GPT5 wrote an edge detection script to see where "golden dog feet" met "bright green grass" to prove to me that there were only 4 legs. The script found 5, and GPT-5 the…

I just tried to get Gemini to produce an image of a dog with 5 legs to test this out, and it really struggled with that. It either made a normal dog, or turned the tail into a weird appendage. Then I asked both Gemini and Grok to count the legs, both kept saying 4. Gemini just refused to consider it was actually wrong. Grok seemed to have an existential crisis when I told it it was wrong, becoming convinced that I ha…

I feel a weird mix of extreme amusement and anger that there's a fleet of absurdly powerful, power-hungry servers sitting somewhere being used to process this problem for 2.5 minutes
Post reply on HN