Live data from Hacker News

Gemini 3 Pro Model Card [pdf]

storage.googleapis.com

331–340 of 359 posts

Re: Gemini 3 Pro Model Card [pdf]

#331
post #69

Earlier quoted context omitted.

what is Google Antigravity?

Couple patterns this could follow Speed? (Flash, Flash-Lite, Antigravity) this is my guess. Bonus: maybe Gemini Diffusion soon? Space? (Google Cloud, Google Antigravity?) Clothes? (A light wearable -> Antigravity?) Gaming? (Ghosting/nontangibility -> antigravity?)

[deleted]

Re: Gemini 3 Pro Model Card [pdf]

#332
post #229

Earlier quoted context omitted.

Google was never really late. Where people perceived Google to have dropped the ball was in its productization of AI. The Google's Bard branding stumble was so (hilariously) bad that it threw a lot of people off the scent. My hunch is that, aside from "safety" reasons, the Google Books lawsuit left some copyright wounds that Google did not want to reopen.

Google’s productization is still rather poor. If I want to use OpenAI’s models, I go to their website, look up the price and pay it. For Google’s, I need to figure out whether I want AI Studio or Google Cloud Code Assist or AI Ultra, etc, and if this is for commercial use where I need to prevent Google from training on my data, figuring out which options work is extra complicated. As of a couple weeks ago (the last t…

Anthropic sign-on is surprisingly bad.

Re: Gemini 3 Pro Model Card [pdf]

#333
post #23

Benchmarks from page 4 of the model card: | Benchmark | 3 Pro | 2.5 Pro | Sonnet 4.5 | GPT-5.1 | |-----------------------|-----------|---------|------------|-----------| | Humanity's Last Exam | 37.5% | 21.6% | 13.7% | 26.5% | | ARC-AGI-2 | 31.1% | 4.9% | 13.6% | 17.6% | | GPQA Diamond | 91.9% | 86.4% | 83.4% | 88.1% | | AIME 2025 | | | | | | (no tools) | 95.0% | 88.0% | 87.0% | 94.0% | | (code execution) | 100% | -…

The vending-bench 2 benchmark is kind of nutty [1].

Not sure 360 days is enough of a sample really but it's an interesting take on AI benchmarks.

Are there any other interesting benchmarks to look at?

[1] https://andonlabs.com/evals/vending-bench-2

Re: Gemini 3 Pro Model Card [pdf]

#335

Earlier quoted context omitted.

Google was catastrophically traumatized throughout the org when they had that photos AI mislabel black people as gorillas. They turned the safety and caution knobs up to 12 after that for years, really until OpenAI came along and ate their lunch.

It still haunts them. Even in the brand-new Gemini-based rework of Photos search and image recognition, "gorilla" is a completely blacklisted word.

It should be blocklisted instead. How insensitive of them.

Re: Gemini 3 Pro Model Card [pdf]

#336
post #75

It is interesting that the Gemini 3 beats every other model on these benchmarks, mostly by a wide margin, but not on SWE Bench. Sonnet is still king here and all three look to be basically on the same level. Kind of wild to see them hit such a wall when it comes to agentic coding

swebench is (1) terrible and (2) saturated

Re: Gemini 3 Pro Model Card [pdf]

#338

Earlier quoted context omitted.

A legally binding pinky swear LOL

with fineprint somewhere on page #67, that there are exceptions.

Who needs fine print when there is an SRE with access to the servers who is friends with a research director who gets paid more if the score goes up?

Re: Gemini 3 Pro Model Card [pdf]

#339
post #324

Earlier quoted context omitted.

Wow. They must have had some major breakthrough. Those scores are truly insane. O_O Models have begun to fairly thoroughly saturate "knowledge" and such, but there are still considerable bumps there But the _big news_, and the demonstration of their achievement here, are the incredible scores they've racked up here for what's necessary for agentic AI to become widely deployable. t2-bench. Visual comprehension. Comput…

SWE-Bench Verified | 76.2% | 59.6% | 77.2% | 76.3% is actually insane.

Anthropomorphic found their corner and are standing strong there.

Re: Gemini 3 Pro Model Card [pdf]

#340
SWE-Bench is disappointing not because it is lower than Claude, but because improving on all other domains of knowledge didn't help. So does this mean that this is actually a MoE model in the sense that one expert doesn't talk to the other ?
Post reply on HN