Live data from Hacker News

Gemini 3 Pro: the frontier of vision AI

blog.google

41–50 of 309 posts

Re: Gemini 3 Pro: the frontier of vision AI

#42
post #9

Interesting "ScreenSpot Pro" results: 72.7% Gemini 3 Pro 11.4% Gemini 2.5 Pro 49.9% Claude Opus 4.5 3.50% GPT-5.1 ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use https://arxiv.org/abs/2504.07981

That is... astronomically different. Is GPT-5.1 downscaling and losing critical information or something? How could it be so different?

Re: Gemini 3 Pro: the frontier of vision AI

#43
post #26

"Gemini 3 Pro represents a generational leap from simple recognition to true visual and spatial reasoning." Prompt: "wine glass full to the brim" Image generated: 2/3 full wine glass. True visual and spatial reasoning denied.

I actually did this prompt and found that it worked with a single nudge on a followup prompt. My first shot got me a wine glass that was almost full but not quite. I told it I wanted it full to the top - another drop would overflow. The second shot was perfectly full.

The correction I expect to give to an intern, not a junior person.

Re: Gemini 3 Pro: the frontier of vision AI

#44

I'm really fascinate by the opportunities to analyze videos. The amount of tokens it compresses down to, and what you can reason across those tokens, is incredible.

The actual token calculations with input videos for Gemini 3 Pro is...confusing.

https://ai.google.dev/gemini-api/docs/media-resolution

Re: Gemini 3 Pro: the frontier of vision AI

#45
post #9

Interesting "ScreenSpot Pro" results: 72.7% Gemini 3 Pro 11.4% Gemini 2.5 Pro 49.9% Claude Opus 4.5 3.50% GPT-5.1 ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use https://arxiv.org/abs/2504.07981

I was surprised at how poorly GPT-5 did in comparison to Opus 4.1 and Gemini 2.5 on a pretty simple OCR task a few months ago - I should run that again against the latest models and see how they do. https://simonwillison.net/2025/Aug/29/the-perils-of-vibe-cod...

Re: Gemini 3 Pro: the frontier of vision AI

#46

The document is paints a super impressive picture, but the core constraint of “network connection to Google required so we can harvest your data” is still a big showstopper for me (and all cloud-based AI tooling, really). I’d be curious to see how well something like this can be distilled down for isolated acceleration on SBCs or consumer kit, because that’s where the billions to be made reside (factories, remote sit…

People with your concerns probably make up 1% of the market if that. Also I don’t upload stuff I’m worried about Google seeing. I wonder if they will allows special plans for corporations

I’m very curious where you get that number from, because I thought the same thing until I got a job inside that market and realized how much more vast it actually is. The revenue numbers might not be as big as Big Tech, but the product market is shockingly vast. My advice is not to confuse Big Tech revenues for total market size, because they bring in such revenue by catering to everyone, rather than specific segments or niches; a McDonald’s will always do more volume than a steakhouse, but it doesn’t mean the market for steakhouses is small enough to ignore.

As for this throwaway line:

> Also I don’t upload stuff I’m worried about Google seeing.

You do realize that these companies harvest even private data, right? Like, even in places you think you own, or that you pay for, they’re mining for revenue opportunities and using you as the product even when you’re a customer, right?

> I wonder if they will allows special plans for corporations

They do, but no matter how much redlining Legal does to protect IP interests, the consensus I keep hearing is “don’t put private or sensitive corporate data into third-parties because no legal agreement will sufficiently protect us from harm if they steal our IP or data”. Just look at the glut of lawsuits against Apple, Google, Microsoft, etc from smaller companies that trusted them to act in good faith but got burned for evidence that you cannot trust these entities.

Re: Gemini 3 Pro: the frontier of vision AI

#47

I do some electrical drafting work for construction and throw basic tasks at LLMs. I gave it a shitty harness and it almost 1 shotted laying out outlets in a room based on a shitty pdf. I think if I gave it better control it could do a huge portion of my coworkers jobs very soon

Can you give an example of the sort of harness you used for that? Would love to play around with it

Re: Gemini 3 Pro: the frontier of vision AI

#48
Well

It is the first model to get partial-credit on an LLM image test I have. Which is counting the legs of a dog. Specifically, a dog with 5 legs. This is a wild test, because LLMs get really pushy and insistent that the dog only has 4 legs.

In fact GPT5 wrote an edge detection script to see where "golden dog feet" met "bright green grass" to prove to me that there were only 4 legs. The script found 5, and GPT-5 then said it was a bug, and adjusted the script sensitivity so it only located 4, lol.

Anyway, Gemini 3, while still being unable to count the legs first try, did identify "male anatomy" (it's own words) also visible in the picture. The 5th leg was approximately where you could expect a well endowed dog to have a "5th leg".

That aside though, I still wouldn't call it particularly impressive.

As a note, Meta's image slicer correctly highlighted all 5 legs without a hitch. Maybe not quite a transformer, but interesting that it could properly interpret "dog leg" and ID them. Also the dog with many legs (I have a few of them) all had there extra legs added by nano-banana.

Re: Gemini 3 Pro: the frontier of vision AI

#49

Well It is the first model to get partial-credit on an LLM image test I have. Which is counting the legs of a dog. Specifically, a dog with 5 legs. This is a wild test, because LLMs get really pushy and insistent that the dog only has 4 legs. In fact GPT5 wrote an edge detection script to see where "golden dog feet" met "bright green grass" to prove to me that there were only 4 legs. The script found 5, and GPT-5 the…

this is hilarious and incredibly interesting at the same time! thanks for writing it up.

Re: Gemini 3 Pro: the frontier of vision AI

#50
post #26

"Gemini 3 Pro represents a generational leap from simple recognition to true visual and spatial reasoning." Prompt: "wine glass full to the brim" Image generated: 2/3 full wine glass. True visual and spatial reasoning denied.

I actually did this prompt and found that it worked with a single nudge on a followup prompt. My first shot got me a wine glass that was almost full but not quite. I told it I wanted it full to the top - another drop would overflow. The second shot was perfectly full.

did it return the exact same glass and surrounding imagery, just with more wine?
Post reply on HN