Live data from Hacker News

Gemini 3 Pro Model Card [pdf]

storage.googleapis.com

341–350 of 359 posts

Re: Gemini 3 Pro Model Card [pdf]

#341
post #16

They scored a 31.1% on ARC AGI 2 which puts them in first place. Also notable which models they include for comparison: Gemini 2.5 Pro, Claude Sonnet 4.5, and GPT-5.1. That seems like a minor snub against Grok 4 / Grok 4.1.

About ARC 2: I would want to hear more detail about prompts, frameworks, thinking time, etc., but they don't matter too much. The main caveat would be that this is probably on the public test set, so could be in pretraining, and there could even be some ARC-focussed post-training - I think we don't know yet and might never know. But for any reasonable setup, if no egregious cheating, that is an amazing score on ARC 2…

This is on the semi-private set

* https://x.com/arcprize/status/1990820655411909018

* https://arcprize.org/guide

Re: Gemini 3 Pro Model Card [pdf]

#342
post #75

It is interesting that the Gemini 3 beats every other model on these benchmarks, mostly by a wide margin, but not on SWE Bench. Sonnet is still king here and all three look to be basically on the same level. Kind of wild to see them hit such a wall when it comes to agentic coding

I don't know if this is true but I believe Anthropic has for a long time illegally used user prompts for training, without user consent.

Re: Gemini 3 Pro Model Card [pdf]

#343

Earlier quoted context omitted.

Google was never really late. Where people perceived Google to have dropped the ball was in its productization of AI. The Google's Bard branding stumble was so (hilariously) bad that it threw a lot of people off the scent. My hunch is that, aside from "safety" reasons, the Google Books lawsuit left some copyright wounds that Google did not want to reopen.

Bard was horrible compared to the competition of the time. Gemini 1.0 was strictly worse than GPT-3.5 and was unusable due to "safety" features. Google followed that up with 1.5 which was still worse than GPT-3.5 and unbelievably far behind GPT-4. At this same time Google had their "black nazi" scandals. With Gemini 2.0 finally had a model that was at least useful for OCR and with their fash series a model that, whil…

Gemini 1.5 Pro was definitely useful at OCR. I used it for that on the free tier.

Re: Gemini 3 Pro Model Card [pdf]

#344
post #288

Earlier quoted context omitted.

From https://lastexam.ai/ : "The dataset consists of 2,500 challenging questions across over a hundred subjects. We publicly release these questions, while maintaining a private test set of held out questions to assess model overfitting ." [emphasis mine] While the private questions don't seem to be included in the performance results, HLE will presumably flag any LLM that appears to have gamed its scores based on th…

You have to trust that the LLM provider isn't copying the questions when Humanities Last Exam runs the test.

There are only eleventy trillion dollars shifting around based on the results, so nobody has any reason to lie.

Re: Gemini 3 Pro Model Card [pdf]

#345
post #249

Earlier quoted context omitted.

Wow. They must have had some major breakthrough. Those scores are truly insane. O_O Models have begun to fairly thoroughly saturate "knowledge" and such, but there are still considerable bumps there But the _big news_, and the demonstration of their achievement here, are the incredible scores they've racked up here for what's necessary for agentic AI to become widely deployable. t2-bench. Visual comprehension. Comput…

The problem is that we know in advance what is the benchmark, so Humanity's Last Exam for example, it's way easier to optimize your model when you have seen the questions before.

not possible on ARC-AGI, AFAIK

Re: Gemini 3 Pro Model Card [pdf]

#346
post #90
post #75

It is interesting that the Gemini 3 beats every other model on these benchmarks, mostly by a wide margin, but not on SWE Bench. Sonnet is still king here and all three look to be basically on the same level. Kind of wild to see them hit such a wall when it comes to agentic coding

This might also hint at SWE struggling to capture what “being good at coding” means. Evals are hard.

It is just Python and Django. It might indicate qualities in other technologies, but it is not very good benchmark.

Re: Gemini 3 Pro Model Card [pdf]

#347
post #75

It is interesting that the Gemini 3 beats every other model on these benchmarks, mostly by a wide margin, but not on SWE Bench. Sonnet is still king here and all three look to be basically on the same level. Kind of wild to see them hit such a wall when it comes to agentic coding

From my personal experience using the CLI agentic coding tools, I think gemini-cli is fairly on par with the rest in terms of the planning/code that is generated. However, when I recently tried qwen-code, it gave me a better sense of reasoning and structure that geimini. Claude definitely has it's own advantages but is expensive(at least for some if not for all). My point is, although the model itself may have perfor…

Yeah, you can see this even by just running claude-code against other models. For example, DeepSeek used as a backend for CC tends to produce results mostly competitive with Sonnet 4.5 A lot is just in the tooling and prompting.

Re: Gemini 3 Pro Model Card [pdf]

#348

Earlier quoted context omitted.

I think Anthropic is reading the room, and just going to go hard on being "the" coding model. I suppose they feel that if they can win that, they can get an ROI without having to do full blown multimodality at the highest level. It's probably pretty liberating, because you can make a "spikey" intelligence with only one spike to really focus on.

Codex has been good enough to me and it’s much cheaper. I code non-trivial stuff with it like multi-threaded code and at least for my style of AI coding which is to do fairly small units of work with multiple revisions it is good enough for me to not to even consider the competition. Just giving you a perspective on how the benchmarks might not be important at all for some people and how Claude may have a difficult t…

My issue with codex is needing to run it in wsl in windows, due to it spamming confirmation requests for running even the safest of commands (eg list directory contents, read file, git status) which in turn adds an extra layer of complexity hooking it up via MCP to anything running in windows outside of wsl (like say figma)

In Claude on the other hand, MCP connections really do seem to ‘just work’

Re: Gemini 3 Pro Model Card [pdf]

#349
post #240

Earlier quoted context omitted.

Can you explain what you mean by this? iPhone was the end of Blackberry. It seems reasonable that a smarter, cheaper, faster model would obsolete anything else. ChatGPT has some brand inertia, but not that much given it's barely 2 years old.

Ask yourself why Microsoft Teams won. These are business tools first and foremost.

That's an odd take. Teams doesn't have the leading market share in videoconferencing, Zoom does. I can't judge what it's like because I've never yet had to use Teams - not a single company that we deal with uses it, it's all Zoom and Chime - but I do hear friends who have to use it complain about it all the time. (Zoom is better than it used to be, but for all that is holy please get rid of the floating menu when we're sharing screens)

Re: Gemini 3 Pro Model Card [pdf]

#350
post #23

Benchmarks from page 4 of the model card: | Benchmark | 3 Pro | 2.5 Pro | Sonnet 4.5 | GPT-5.1 | |-----------------------|-----------|---------|------------|-----------| | Humanity's Last Exam | 37.5% | 21.6% | 13.7% | 26.5% | | ARC-AGI-2 | 31.1% | 4.9% | 13.6% | 17.6% | | GPQA Diamond | 91.9% | 86.4% | 83.4% | 88.1% | | AIME 2025 | | | | | | (no tools) | 95.0% | 88.0% | 87.0% | 94.0% | | (code execution) | 100% | -…

These numbers are impressive, at least to say. It looks like Google has produced a beast that will raise the bar even higher. What's even more impressive is how Google came into this game late and went from producing a few flops to being the leader at this (actually, they already achieved the title with 2.5 Pro). What makes me even more curious is the following > Model dependencies: This model is not a modification o…

There are no leaders. Every other month a new LLM model comes out and it outperforms the previous ones by a small margin, the benchmarks always look good (probably because the models are trained on the answers) but then in practice they are basically indistinguishable from the previous ones (take GPT4 vs 5). We've been in this loop since around the release of ChatGPT 4 where all the main players started this cycle.

The biggest strides in the last 6-8 months have been in generative AIs, specifically for animation.

Post reply on HN