Live data from Hacker News

Gemini 3 Pro Model Card [pdf]

storage.googleapis.com

131–140 of 359 posts

Re: Gemini 3 Pro Model Card [pdf]

#132

I know this is a little controversial but the lack of performance on SWE-bench is hugely disappointing I think economically. These models don’t have any viable path to profitability if they can’t take engineering jobs.

I thought that but it does do a lot better on other benchmarks. Perhaps SWE bench just doesn't capture a lot of the improvement? If the web design improvements people have been posting on twitter, I suspect this will be a huge boon for developers. SWE benchmark is really testing bugfixing/feature dev more. Anyway let's see. I'm still hyped!

That would be great! But AI is a bubble if these models can’t do serious engineering work.

Re: Gemini 3 Pro Model Card [pdf]

#133
post #23

Benchmarks from page 4 of the model card: | Benchmark | 3 Pro | 2.5 Pro | Sonnet 4.5 | GPT-5.1 | |-----------------------|-----------|---------|------------|-----------| | Humanity's Last Exam | 37.5% | 21.6% | 13.7% | 26.5% | | ARC-AGI-2 | 31.1% | 4.9% | 13.6% | 17.6% | | GPQA Diamond | 91.9% | 86.4% | 83.4% | 88.1% | | AIME 2025 | | | | | | (no tools) | 95.0% | 88.0% | 87.0% | 94.0% | | (code execution) | 100% | -…

That looks impressive, but some of the are a bit out of date. On Terminal-Bench 2 for example, the leader is currently "Codex CLI (GPT-5.1-Codex)" at 57.8%, beating this new release.

That's a different model not in the chart. They're not going to include hundreds of fine tunes in a chart like this.

Re: Gemini 3 Pro Model Card [pdf]

#134

Curiously, this website seems to be blocked in Spain for whatever reason, and the website's certificate is served by `allot.com/emailAddress=info@allot.com` which obviously fails... Anyone happen to know why? Is this website by any change sharing information on safe medical abortions or women's rights, something which has gotten websites blocked here before?

Creator of pixeldrain here. I have no idea why my site is blocked in Spain, but it's a long running issue. I actually never discovered who was responsible for the blockade, until I read this comment. I'm going to look into Allot and send them an email. EDIT: Also, your DNS provider is censoring (and probably monitoring) your internet traffic. I would switch to a different provider.

> EDIT: Also, your DNS provider is censoring (and probably monitoring) your internet traffic. I would switch to a different provider.

Yeah, that was via my ISPs DNS resolver (Vodafone), switching the resolver works :)

The responsible party is ultimately our government who've decided it's legal to block a wide range of servers and websites because some people like to watch illegal football streams. I think Allot is just the provider of the technology.

Re: Gemini 3 Pro Model Card [pdf]

#135

Earlier quoted context omitted.

> Anyone with money can trivially catch up to a state of the art model from six months ago. How come apple is struggling then?

It looks more like a strategic decision tbh. The may want to use 3rd party or just wait for AI to be more stable to see how people actually use it instead of adding slop in the core of their product.

In contrast to Microsoft, who puts Copilot buttons everywhere and succeeds only in annoying their customers.

Re: Gemini 3 Pro Model Card [pdf]

#136
post #23

Benchmarks from page 4 of the model card: | Benchmark | 3 Pro | 2.5 Pro | Sonnet 4.5 | GPT-5.1 | |-----------------------|-----------|---------|------------|-----------| | Humanity's Last Exam | 37.5% | 21.6% | 13.7% | 26.5% | | ARC-AGI-2 | 31.1% | 4.9% | 13.6% | 17.6% | | GPQA Diamond | 91.9% | 86.4% | 83.4% | 88.1% | | AIME 2025 | | | | | | (no tools) | 95.0% | 88.0% | 87.0% | 94.0% | | (code execution) | 100% | -…

We knew it would be a big jump and while it certainly is in many areas - its definitely not "groundbreaking/huge leap" worthy like some were thinking from looking at these numbers.

I feel like many will be pretty disappointed by their self created expectations for this model when they end up actually using it and it turns out to be fairly similar to other frontier models.

Personally I'm very interested in how they end up pricing it.

Re: Gemini 3 Pro Model Card [pdf]

#137
post #86

Earlier quoted context omitted.

I care very little about model personality outside of sycophancy. The thing about gemini is that it's notorious for its low self esteem. Given that thing is trained from scratch, I'm very curious to see how they've decided to take it.

given how often these llms are wrong, doesnt it make sense that they are less confident?

Indeed. But I've had experiences with gemini-2.5-pro-exp where its thoughts could be described as "rejected from the prom" vibes. It's not like I abused it either, it was running into loops because it was unable to properly patch a file.

Re: Gemini 3 Pro Model Card [pdf]

#138

Curiously, this website seems to be blocked in Spain for whatever reason, and the website's certificate is served by `allot.com/emailAddress=info@allot.com` which obviously fails... Anyone happen to know why? Is this website by any change sharing information on safe medical abortions or women's rights, something which has gotten websites blocked here before?

do you know about the cloudflare and laliga issues? might be that

Was my first instinct, went looking if there was any games being played today but seems not, so unlikely to be the cause.

Re: Gemini 3 Pro Model Card [pdf]

#139
post #98

Earlier quoted context omitted.

These numbers are impressive, at least to say. It looks like Google has produced a beast that will raise the bar even higher. What's even more impressive is how Google came into this game late and went from producing a few flops to being the leader at this (actually, they already achieved the title with 2.5 Pro). What makes me even more curious is the following > Model dependencies: This model is not a modification o…

At least at the moment, coming in late seems to matter little. Anyone with money can trivially catch up to a state of the art model from six months ago. And as others have said, late is really a function of spigot, guardrails, branding, and ux, as much as it is being a laggard under the hood.

Being known as a company that is always six months late than the competitors isn't something to brag about...

Re: Gemini 3 Pro Model Card [pdf]

#140

Earlier quoted context omitted.

Not apples-to-apples. "Codex CLI (GPT-5.1-Codex)", which the site refers to, adds a specific agentic harness, whereas the Gemini 3 Pro seems to be on a standard eval harness. It would be interesting to see the apples-to-apples figure, i.e. with Google's best harness alongside Codex CLI.

Do you mean that Gemini 3 Pro is "vanilla" like GPT 5.1 (non-Codex)?

Yes, two things: 1. GPT-5.1 Codex is a fine tune, not the "vanilla" 5.1 2. More importantly, GPT 5.1 Codex achieves its performance when used with a specific tool (Codex CLI) that is optimized for GPT 5.1 Codex. But when labs evaluate the models, they have to use a standard tool to make the comparisons apples-to-apples.

Will be interesting to see what Google releases that's coding-specific to follow Gemini 3.

Post reply on HN