Live data from Hacker News

Gemini 3 Pro Model Card [pdf]

storage.googleapis.com

81–90 of 359 posts

Re: Gemini 3 Pro Model Card [pdf]

#81
post #37
post #13

Earlier quoted context omitted.

Why? These models just leapfrog each other as time advances. One month Gemini is on top, then ChatGPT, then Anthropic. Not sure why everyone gets FOMO whenever a new version gets released.

Considering GPT 5 was only recently released, it's very unlikely GPT will achieve these scores in just a couple of months. If they had something this good in the oven, they'd probably left the GPT 5 name to it. Or maybe Google just benchmaxxed and this doesn't translate at all in real world performance.

They do have unreleased Olympiad Gold-winning models that are definitely better than GPT5.

TBD if that performance generalizes to other real world tasks.

Re: Gemini 3 Pro Model Card [pdf]

#82

Earlier quoted context omitted.

These numbers are impressive, at least to say. It looks like Google has produced a beast that will raise the bar even higher. What's even more impressive is how Google came into this game late and went from producing a few flops to being the leader at this (actually, they already achieved the title with 2.5 Pro). What makes me even more curious is the following > Model dependencies: This model is not a modification o…

Google was never really late. Where people perceived Google to have dropped the ball was in its productization of AI. The Google's Bard branding stumble was so (hilariously) bad that it threw a lot of people off the scent. My hunch is that, aside from "safety" reasons, the Google Books lawsuit left some copyright wounds that Google did not want to reopen.

Oh, I remember the times when I compared Gemini with ChatGPT and Claude. Gemini was so far behind, it was barely usable. And now they are pushing the boundries.

Re: Gemini 3 Pro Model Card [pdf]

#85
post #77

Earlier quoted context omitted.

These numbers are impressive, at least to say. It looks like Google has produced a beast that will raise the bar even higher. What's even more impressive is how Google came into this game late and went from producing a few flops to being the leader at this (actually, they already achieved the title with 2.5 Pro). What makes me even more curious is the following > Model dependencies: This model is not a modification o…

What does it mean nowadays to start from scratch? At least in the open scene, most of the post-training data is generated by other LLMs.

They had to start with a base model, that part I am certain of

Re: Gemini 3 Pro Model Card [pdf]

#86

There needs to be a sycophancy benchmark in these comparisons. More baseless praise and false agreement = lower score.

I care very little about model personality outside of sycophancy. The thing about gemini is that it's notorious for its low self esteem. Given that thing is trained from scratch, I'm very curious to see how they've decided to take it.

given how often these llms are wrong, doesnt it make sense that they are less confident?

Re: Gemini 3 Pro Model Card [pdf]

#87
post #27
post #13

Earlier quoted context omitted.

Why? These models just leapfrog each other as time advances. One month Gemini is on top, then ChatGPT, then Anthropic. Not sure why everyone gets FOMO whenever a new version gets released.

I think google is uniquely well placed to make a profitable business out of AI: They make their own TPUs so don't have to pay ridiculous amounts of money to Nvidia, they have a great depth of talent in building models, they've got loads of data they can use for training and they've got a huge existing customer base who can buy their AI offerings. I don't think any other company has all these ingredients.

The bear case for Google was always the business side would cannibalize the AI side. AI makes search redundant which kills the golden goose

Re: Gemini 3 Pro Model Card [pdf]

#88
post #67
post #37

Earlier quoted context omitted.

Considering GPT 5 was only recently released, it's very unlikely GPT will achieve these scores in just a couple of months. If they had something this good in the oven, they'd probably left the GPT 5 name to it. Or maybe Google just benchmaxxed and this doesn't translate at all in real world performance.

GPT 5 was released more than 3 months ago. Gemini 2.5 was released less than 8 months ago.

If not this model, Google at some point is going to get and stay ahead just because they have so many more people and compute resources they can throw at many directions while the others have to make the right choices with how they use their resources each time. Took a while to channel their numbers into a product direction but now I don't think they're going to let up

Re: Gemini 3 Pro Model Card [pdf]

#89
post #69

Title of the document is "[Gemini 3 Pro] External Model Card - November 18, 2025 - v2", in case you needed further confirmation that the model will be released today. Also interesting to know that Google Antigravity (antigravity.google / https://github.com/Google-Antigravity ?) leaked. I remember seeing this subdomain recently. Probably Gemini 3 related as well. Org was created on 2025-11-04T19:28:13Z ( https://api.g…

what is Google Antigravity?

The ASI figured out zero point energy from first principles

Re: Gemini 3 Pro Model Card [pdf]

#90
post #75

It is interesting that the Gemini 3 beats every other model on these benchmarks, mostly by a wide margin, but not on SWE Bench. Sonnet is still king here and all three look to be basically on the same level. Kind of wild to see them hit such a wall when it comes to agentic coding

This might also hint at SWE struggling to capture what “being good at coding” means.

Evals are hard.

Post reply on HN