Live data from Hacker News

Gemini 3 Pro Model Card [pdf]

storage.googleapis.com

181–190 of 359 posts

Re: Gemini 3 Pro Model Card [pdf]

#181
post #23

Benchmarks from page 4 of the model card: | Benchmark | 3 Pro | 2.5 Pro | Sonnet 4.5 | GPT-5.1 | |-----------------------|-----------|---------|------------|-----------| | Humanity's Last Exam | 37.5% | 21.6% | 13.7% | 26.5% | | ARC-AGI-2 | 31.1% | 4.9% | 13.6% | 17.6% | | GPQA Diamond | 91.9% | 86.4% | 83.4% | 88.1% | | AIME 2025 | | | | | | (no tools) | 95.0% | 88.0% | 87.0% | 94.0% | | (code execution) | 100% | -…

Should I assume the GPT-5.1 it is compared against is the pro version?

Re: Gemini 3 Pro Model Card [pdf]

#182

Earlier quoted context omitted.

> EDIT: Also, your DNS provider is censoring (and probably monitoring) your internet traffic. I would switch to a different provider. Yeah, that was via my ISPs DNS resolver (Vodafone), switching the resolver works :) The responsible party is ultimately our government who've decided it's legal to block a wide range of servers and websites because some people like to watch illegal football streams. I think Allot is ju…

My site has nothing to do with football though. And Allot seems to be running the DNS server that your ISP uses so they are directly responsible for the block.

La Liga (the football company) likes to send out takedown notices to anyone who may host anything that looks like a football to protect their precious games, no matter the collateral damage or the lack of any requirements to show damage. They have the right to block anything in Spain at their discretion either by DNS or IP. They do seem to work in good faith if you talk to them, though, and if you can either remove sites or content when they ask.

Re: Gemini 3 Pro Model Card [pdf]

#183
post #23

Benchmarks from page 4 of the model card: | Benchmark | 3 Pro | 2.5 Pro | Sonnet 4.5 | GPT-5.1 | |-----------------------|-----------|---------|------------|-----------| | Humanity's Last Exam | 37.5% | 21.6% | 13.7% | 26.5% | | ARC-AGI-2 | 31.1% | 4.9% | 13.6% | 17.6% | | GPQA Diamond | 91.9% | 86.4% | 83.4% | 88.1% | | AIME 2025 | | | | | | (no tools) | 95.0% | 88.0% | 87.0% | 94.0% | | (code execution) | 100% | -…

Wow. They must have had some major breakthrough. Those scores are truly insane. O_O

Models have begun to fairly thoroughly saturate "knowledge" and such, but there are still considerable bumps there

But the _big news_, and the demonstration of their achievement here, are the incredible scores they've racked up here for what's necessary for agentic AI to become widely deployable. t2-bench. Visual comprehension. Computer use. Vending-Bench. The sorts of things that are necessary for AI to move beyond an auto-researching tool, and into the realm where it can actually handle complex tasks in the way that businesses need in order to reap rewards from deploying AI tech.

Will be very interesting to see what papers are published as a result of this, as they have _clearly_ tapped into some new avenues for training models.

And here I was, all wowed, after playing with Grok 4.1 for the past few hours! xD

Re: Gemini 3 Pro Model Card [pdf]

#184

Earlier quoted context omitted.

On reddit I see it's already available on cursor https://www.reddit.com/r/Bard/comments/1p093fb/gemini_3_in_c...

Interesting, it doesn't show up for me in Cursor yet.

you need to manually add the custom model gemini-3-pro-preview

Re: Gemini 3 Pro Model Card [pdf]

#185
post #69

Title of the document is "[Gemini 3 Pro] External Model Card - November 18, 2025 - v2", in case you needed further confirmation that the model will be released today. Also interesting to know that Google Antigravity (antigravity.google / https://github.com/Google-Antigravity ?) leaked. I remember seeing this subdomain recently. Probably Gemini 3 related as well. Org was created on 2025-11-04T19:28:13Z ( https://api.g…

what is Google Antigravity?

possibly https://xkcd.com/353/

Re: Gemini 3 Pro Model Card [pdf]

#186
post #133

Earlier quoted context omitted.

That looks impressive, but some of the are a bit out of date. On Terminal-Bench 2 for example, the leader is currently "Codex CLI (GPT-5.1-Codex)" at 57.8%, beating this new release.

That's a different model not in the chart. They're not going to include hundreds of fine tunes in a chart like this.

It's not just one of many fine tunes; it's the default model used by OpenAI's official tools.

Re: Gemini 3 Pro Model Card [pdf]

#187
> TPUs are specifically designed to handle the massive computations involved in training LLMs and can speed up training considerably compared to CPUs.

That seems like a low bar. Who's training frontier LLMs on CPUs? Surely they meant to compare TPUs to GPUs. If "this is faster than a CPU for massively parallel AI training" is the best you can say about it, that's not very impressive.

Re: Gemini 3 Pro Model Card [pdf]

#188

Earlier quoted context omitted.

> I guess improvements will be incremental from here on out. What do you mean? These coding leaderboards were at single digits about a year ago and are now in the seventies. These frontier models are arguably already better at the benchmark that any single human - it's unlikely that any particular human dev is knowledgeable to tackle the full range of diverse tasks even in the smaller SWE-Bench Verified within a reas…

A new benchmark comes out, it's designed so nothing does well at it, the models max it out, and the cycle repeats. This could either describe massive growth of LLM coding abilities or a disconnect between what the new benchmarks are measuring & why new models are scoring well after enough time. In the former assumption there is no limit to the growth of scores... but there is also not very much actual growth (if any…

Your mileage may vary, but for me, working today with the latest version of Claude Code on a non-trivial python web dev project, I do absolutely feel that I can hand over to the AI coding tasks that are 10 times more complex or time consuming than what I could hand over to copilot or windsurf a year ago. It's still nowhere close to replacing me, but I feel that I can work at a significantly higher level.

What field are you in where you feel that there might not have been any growth in capabilities at all?

EDIT: Typo

Re: Gemini 3 Pro Model Card [pdf]

#189

So does google actually have a claude console alternative currently?

Noteworthily, although Gemini 3 Pro seems to have much benchmark scores than other models across the board (including compared to Claude), it's not the case for coding, where it appears to score essentially the same as the others. I wonder why that is. So far, IMHO, Claude Code remains significantly better than Gemini CLI. We'll see whether that changes with Gemini 3.

> I wonder why that is.

That's because coding is currently the only reliable benchmark where reasoning capabilities transfer to predict capabilities for other professions like law. Coding is the only area where they are shy to release numbers. All these exam scores are fakeable by gaming those benchmarks.

Re: Gemini 3 Pro Model Card [pdf]

#190

> TPUs are specifically designed to handle the massive computations involved in training LLMs and can speed up training considerably compared to CPUs. That seems like a low bar. Who's training frontier LLMs on CPUs? Surely they meant to compare TPUs to GPUs. If "this is faster than a CPU for massively parallel AI training" is the best you can say about it, that's not very impressive.

It's a typo
Post reply on HN