Live data from Hacker News

Gemini 3 Pro: the frontier of vision AI

blog.google

261–270 of 309 posts

Re: Gemini 3 Pro: the frontier of vision AI

#261
post #253

Earlier quoted context omitted.

Again, think about how the models work. They generate text sequentially. Think about how you solve the maze in your mind. Do you draw a line direct to the finish? No, it would be impossible to know what the path was until you had done it. But at that point you have now backtracked several times. So, what could a model _possibly_ be able to do for this puzzle which is "fair game" as a valid solution, other than magica…

> So, what could a model _possibly_ be able to do for this puzzle which is "fair game" as a valid solution, other than magically know an answer by pulling it out of thin air? Represent the maze as a sequence of movements which either continue or end up being forced to backtrack. Basically it would represent the maze as a graph and do a depth-first search, keeping track of what nodes it as visited in its reasoning tok…

And my question to you is “why is that substantially different than writing the correct algorithm to do it”? Im arguing its a myopic view of what we are going to call “intelligence”. And it ignores how human thought works in the same way by using abstractions to move to the next level of reasoning.

In my opinion, being able to write the code to do the thing is effectively the same exact thing as doing the thing in terms of judging if its “able to do” that thing. Its functionality equivalent for evaluating what the “state of the art” is, and honestly is naive to what these models even are. If the model hid the tool calling in the background instead, and only showed you its answer would we say its more intelligent? Because that’s essentially how a lot of these things work already. Because again, the actual “model” is just a text autocomplete engine and it generates from left to right.

Re: Gemini 3 Pro: the frontier of vision AI

#262

Earlier quoted context omitted.

This is a really interesting "data flywheel" -- better model >> more usable data >> even better model

surely there's an upper limit to this though with models literally eating themselves.

Not always, you can improve the loop by putting something real inside, like, a code execution tool, a search engine, a human, other AIs or an API. As long as the model can make use of that external environment its data can improve. By the same logic a human isolated from other humans for a long time might also be in a situation of going crazy.

Practical example - using LLMs to create deep research reports. It pulls over 500 sources into a complex analysis, and after all that compiling and contrasting it generates an article with references, like a wiki page. That text is probably superior to most of its sources in quality. It does not trust any one source completely, it does not even pretend to present the truth, it only summarizes the distribution of information it found on the topic. Imagine scaling wikipedia 1000x by deep-reporting every conceivable topic.

Re: Gemini 3 Pro: the frontier of vision AI

#263

Earlier quoted context omitted.

So the idea is what? What's the successful outcome look like for this test, in your mind? What should good software do? Respond and say there are 5 legs? Or question what kind of dog this even is? Or get confused by a nonsensical picture that doesn't quite match the prompt in a confusing way? Should it understand the concept of a dog and be able to tell you that this isn't a real dog?

You know, I had a potential hire last week, and I was interviewing this one guy whose resume was really strong, it was exceptional in many ways plus his open-source code was looking really tight. But at the beginning of the interview, I always show the candidates the same silly code example with signed integer overflow undefined behavior baked in. I did the same here and asked him if he sees anything unusual with it,…

Does the ability to verbally detect gotchas in short conversations dealing only with text on a screen or white board really map to stronger candidates?

In actual situations you have documentation, editor, tooling, tests, and are a tad less distracted than when dealing with a job interview and all the attendant stress. Isn't the fact that he actually produces quality code in real life a stronger signal of quality?

Re: Gemini 3 Pro: the frontier of vision AI

#264

Earlier quoted context omitted.

That just seems like an arbitrary limitation. Its like asking someone to do answer a math calculation but "no thinking allowed". Like, I guess we can gauge if a model just _knows all knowable things in the universe_ using that method... but anything of any value that you are gauging in terms of 'intelligence', is going to actually be validating their ability to go "outside the scope" of what they actually are (an aut…

By your analogy, the developers of stockfish are better chess players than any grandmaster. Tool use can be a sign of intelligence, but "being able to use a tool to solve a problem" is not the same as "being intelligent enough to solve a specific class of problems".

Im not talking about this being the "best maze solver" and "better at solving mazes than humans". Im saying the model is "intelligent enough" to solve a maze.

And what Im really saying is that we need to stop moving the goal post on what "intelligence" is for these models, and start moving the goal post on what "intelligence" actually _is_. The models are giving us an existential crisis on not only what it might mean to _be_ intelligent, but also how it might actually work in our own brains. Im not saying the current models are skynet, but Im saying I think theres going to be a lot learned by reverse engineering the current generation of models to really dig into how they are encoding things internally.

Re: Gemini 3 Pro: the frontier of vision AI

#265

Well It is the first model to get partial-credit on an LLM image test I have. Which is counting the legs of a dog. Specifically, a dog with 5 legs. This is a wild test, because LLMs get really pushy and insistent that the dog only has 4 legs. In fact GPT5 wrote an edge detection script to see where "golden dog feet" met "bright green grass" to prove to me that there were only 4 legs. The script found 5, and GPT-5 the…

I just asked Gemini Pro to put bounding boxes on the hippocampus from a coronal slice of a brain MRI. Complete fail. There has to be thousands of pictures of coronal brain slices with hippocampal labels out there, but apparently it learned none of it...unless I am doing it wrong.

https://i.imgur.com/1XxYoYN.png

Re: Gemini 3 Pro: the frontier of vision AI

#266

Well It is the first model to get partial-credit on an LLM image test I have. Which is counting the legs of a dog. Specifically, a dog with 5 legs. This is a wild test, because LLMs get really pushy and insistent that the dog only has 4 legs. In fact GPT5 wrote an edge detection script to see where "golden dog feet" met "bright green grass" to prove to me that there were only 4 legs. The script found 5, and GPT-5 the…

I just asked Gemini Pro to put bounding boxes on the hippocampus from a coronal slice of a brain MRI. Complete fail. There has to be thousands of pictures of coronal brain slices with hippocampal labels out there, but apparently it learned none of it...unless I am doing it wrong. https://i.imgur.com/1XxYoYN.png

asked nanobanana to paint the hippocampus red...better, but not close to good. https://imgur.com/a/clwNg1h

Re: Gemini 3 Pro: the frontier of vision AI

#267

Earlier quoted context omitted.

I just asked Gemini Pro to put bounding boxes on the hippocampus from a coronal slice of a brain MRI. Complete fail. There has to be thousands of pictures of coronal brain slices with hippocampal labels out there, but apparently it learned none of it...unless I am doing it wrong. https://i.imgur.com/1XxYoYN.png

asked nanobanana to paint the hippocampus red...better, but not close to good. https://imgur.com/a/clwNg1h

I was a little hopeful when I tried again, but it really seems that it didn't know what it looks like. Maybe few shot with examples?

https://gemini.google.com/share/137812b95b5e

Re: Gemini 3 Pro: the frontier of vision AI

#269

Earlier quoted context omitted.

I just tried to get Gemini to produce an image of a dog with 5 legs to test this out, and it really struggled with that. It either made a normal dog, or turned the tail into a weird appendage. Then I asked both Gemini and Grok to count the legs, both kept saying 4. Gemini just refused to consider it was actually wrong. Grok seemed to have an existential crisis when I told it it was wrong, becoming convinced that I ha…

An interesting test in this vein that I read about in a comment on here is generating a 13 hour clock—I tried just about every prompting trick and clever strategy I could come up with across many image models with no success. I think there's so much training data of 12 hour clocks that just clobbers the instructions entirely. It'll make a regular clock that skips from 11 to 13, or a regular clock with a plaque saying…

[deleted]

Re: Gemini 3 Pro: the frontier of vision AI

#270

I do some electrical drafting work for construction and throw basic tasks at LLMs. I gave it a shitty harness and it almost 1 shotted laying out outlets in a room based on a shitty pdf. I think if I gave it better control it could do a huge portion of my coworkers jobs very soon

I just can't imagine we are close to letting LLMs do electrical work.

What I notice that I don't see talked about much is how "steerable" the output is.

I think this is a big reason 1 shots are used as examples.

Once you get past 1 shots, so much of the output is dependent on the context the previous prompts have created.

Instead of 1 shots , try something that requires 3 different prompts on a subject with uncertainty involved. Do 4 or 5 iterations and often you will get wildly different results.

It doesn't seem like we have a word for this. A "hallucination" is when we know what the output should be and it is just wrong. This is like the user steers the model towards an answer but there is a lot of uncertainty in what the right answer even would be.

To me this always comes back to the problem that the models are not grounded in reality.

Letting LLMs do electric work without grounding in reality would be insane. No pun intended.

Post reply on HN