Live data from Hacker News

Gemini 3 Pro Model Card [pdf]

storage.googleapis.com

261–270 of 359 posts

Re: Gemini 3 Pro Model Card [pdf]

#261
post #23

Benchmarks from page 4 of the model card: | Benchmark | 3 Pro | 2.5 Pro | Sonnet 4.5 | GPT-5.1 | |-----------------------|-----------|---------|------------|-----------| | Humanity's Last Exam | 37.5% | 21.6% | 13.7% | 26.5% | | ARC-AGI-2 | 31.1% | 4.9% | 13.6% | 17.6% | | GPQA Diamond | 91.9% | 86.4% | 83.4% | 88.1% | | AIME 2025 | | | | | | (no tools) | 95.0% | 88.0% | 87.0% | 94.0% | | (code execution) | 100% | -…

Used an AI to populate some of 5.1 thinking's results. Benchmark | Gemini 3 Pro | Gemini 2.5 Pro | Claude Sonnet 4.5 | GPT-5.1 | GPT-5.1 Thinking ---------------------------|--------------|----------------|-------------------|---------|------------------ Humanity's Last Exam | 37.5% | 21.6% | 13.7% | 26.5% | 52% ARC-AGI-2 | 31.1% | 4.9% | 13.6% | 17.6% | 28% GPQA Diamond | 91.9% | 86.4% | 83.4% | 88.1% | 61% AIM 2025…

What? The 4.5 and 5.1 columns aren't thinking in Google's report?

That's a scandal, IMO.

Given that Gemini-3 seems to do "fine" against the thinking versions why didn't they post those results? I get that PMs like to make a splash but that's shockingly dishonest.

Re: Gemini 3 Pro Model Card [pdf]

#262

Earlier quoted context omitted.

A new benchmark comes out, it's designed so nothing does well at it, the models max it out, and the cycle repeats. This could either describe massive growth of LLM coding abilities or a disconnect between what the new benchmarks are measuring & why new models are scoring well after enough time. In the former assumption there is no limit to the growth of scores... but there is also not very much actual growth (if any…

Your mileage may vary, but for me, working today with the latest version of Claude Code on a non-trivial python web dev project, I do absolutely feel that I can hand over to the AI coding tasks that are 10 times more complex or time consuming than what I could hand over to copilot or windsurf a year ago. It's still nowhere close to replacing me, but I feel that I can work at a significantly higher level. What field a…

Claude 3.5 came out in June of last year, and it is imo marginally worse than the AI models currently available for coding. I do not think models are 10x better than 1 year ago, that seems extremely hyperbolic or you are working in a super niche area where that is true.

Re: Gemini 3 Pro Model Card [pdf]

#263

Curiously, this website seems to be blocked in Spain for whatever reason, and the website's certificate is served by `allot.com/emailAddress=info@allot.com` which obviously fails... Anyone happen to know why? Is this website by any change sharing information on safe medical abortions or women's rights, something which has gotten websites blocked here before?

loads fine on Vodafone for me

Re: Gemini 3 Pro Model Card [pdf]

#264
post #70

Earlier quoted context omitted.

Looks like it will be on par with the contenders when it comes to coding. I guess improvements will be incremental from here on out.

> I guess improvements will be incremental from here on out. What do you mean? These coding leaderboards were at single digits about a year ago and are now in the seventies. These frontier models are arguably already better at the benchmark that any single human - it's unlikely that any particular human dev is knowledgeable to tackle the full range of diverse tasks even in the smaller SWE-Bench Verified within a reas…

Google has had a lot of time to optimise for those benchmarks, and just barely made SOTA (or not even SOTA) now. How is that not incremental?

Re: Gemini 3 Pro Model Card [pdf]

#266
post #70

Earlier quoted context omitted.

Looks like it will be on par with the contenders when it comes to coding. I guess improvements will be incremental from here on out.

> I guess improvements will be incremental from here on out. What do you mean? These coding leaderboards were at single digits about a year ago and are now in the seventies. These frontier models are arguably already better at the benchmark that any single human - it's unlikely that any particular human dev is knowledgeable to tackle the full range of diverse tasks even in the smaller SWE-Bench Verified within a reas…

If we're being completely honest, a benchmark is like an honest exam: any set of questions can only be used once when it comes out. Otherwise you're only testing how well people can acquire and memorize exact questions.

Re: Gemini 3 Pro Model Card [pdf]

#267
post #98

Earlier quoted context omitted.

At least at the moment, coming in late seems to matter little. Anyone with money can trivially catch up to a state of the art model from six months ago. And as others have said, late is really a function of spigot, guardrails, branding, and ux, as much as it is being a laggard under the hood.

> Anyone with money can trivially catch up to a state of the art model from six months ago. How come apple is struggling then?

Apple is struggling with _productizing_ LLMs for the mass market, which is a separate task from training a frontier LLM.

To be fair to Apple, so far the only mass market LLM use case so far is just a simple chatbot, and they don't seem to be interested in that. It remains to be seen if what Apple wants to do ("private" LLMs with access to your personal context acting as intimate personal assistants) is even possible to do reliably. It sounds useful, and I do believe it will eventually be possible, but no one is there yet.

They did botch the launch by announcing the Apple Intelligence features before they are ready though.

Re: Gemini 3 Pro Model Card [pdf]

#268
post #229

Earlier quoted context omitted.

Google’s productization is still rather poor. If I want to use OpenAI’s models, I go to their website, look up the price and pay it. For Google’s, I need to figure out whether I want AI Studio or Google Cloud Code Assist or AI Ultra, etc, and if this is for commercial use where I need to prevent Google from training on my data, figuring out which options work is extra complicated. As of a couple weeks ago (the last t…

Not to mention no macOS app. This is probably unimportant to many in the hn audience, but more broadly it matters for your average knowledge worker.

And a REALLY good macOS app.

Like, kind of unreasonably good. You’d expect some perfunctory Electronic app that just barely wraps the website. But no, you get something that feels incredibly polished…more so than a lot of recent apps from Apple…and has powerful integrations into other apps, including text editors and terminals.

Re: Gemini 3 Pro Model Card [pdf]

#269

> Gemini 3 Pro was trained using Google’s Tensor Processing Units (TPUs) NVDA is down 3.26%

If it’s because of that, then honestly it’s as insane as the deepseek thing where all the info was released weeks before but the markt got nervous only when they released an app. I mean info about Gemini 3 is out quite a while now and of course they trained it using TPUs, I didn’t even think that was in question.

I didn't know they only used TPUs.

Re: Gemini 3 Pro Model Card [pdf]

#270

Earlier quoted context omitted.

> EDIT: Also, your DNS provider is censoring (and probably monitoring) your internet traffic. I would switch to a different provider. Yeah, that was via my ISPs DNS resolver (Vodafone), switching the resolver works :) The responsible party is ultimately our government who've decided it's legal to block a wide range of servers and websites because some people like to watch illegal football streams. I think Allot is ju…

My site has nothing to do with football though. And Allot seems to be running the DNS server that your ISP uses so they are directly responsible for the block.

The Spanish courts have allowed la Liga to completely ban every website served by cloudflare during days where there are matches. All Spanish ISPs have to do dns blocking to comply.
Post reply on HN