They scored a 31.1% on ARC AGI 2 which puts them in first place. Also notable which models they include for comparison: Gemini 2.5 Pro, Claude Sonnet 4.5, and GPT-5.1. That seems like a minor snub against Grok 4 / Grok 4.1.
About ARC 2: I would want to hear more detail about prompts, frameworks, thinking time, etc., but they don't matter too much. The main caveat would be that this is probably on the public test set, so could be in pretraining, and there could even be some ARC-focussed post-training - I think we don't know yet and might never know. But for any reasonable setup, if no egregious cheating, that is an amazing score on ARC 2…
Gemini 3 Pro Model Card [pdf]
341–350 of 359 posts
Re: Gemini 3 Pro Model Card [pdf]
#342It is interesting that the Gemini 3 beats every other model on these benchmarks, mostly by a wide margin, but not on SWE Bench. Sonnet is still king here and all three look to be basically on the same level. Kind of wild to see them hit such a wall when it comes to agentic coding
Re: Gemini 3 Pro Model Card [pdf]
#343Earlier quoted context omitted.
Google was never really late. Where people perceived Google to have dropped the ball was in its productization of AI. The Google's Bard branding stumble was so (hilariously) bad that it threw a lot of people off the scent. My hunch is that, aside from "safety" reasons, the Google Books lawsuit left some copyright wounds that Google did not want to reopen.
Bard was horrible compared to the competition of the time. Gemini 1.0 was strictly worse than GPT-3.5 and was unusable due to "safety" features. Google followed that up with 1.5 which was still worse than GPT-3.5 and unbelievably far behind GPT-4. At this same time Google had their "black nazi" scandals. With Gemini 2.0 finally had a model that was at least useful for OCR and with their fash series a model that, whil…
Re: Gemini 3 Pro Model Card [pdf]
#344Earlier quoted context omitted.
From https://lastexam.ai/ : "The dataset consists of 2,500 challenging questions across over a hundred subjects. We publicly release these questions, while maintaining a private test set of held out questions to assess model overfitting ." [emphasis mine] While the private questions don't seem to be included in the performance results, HLE will presumably flag any LLM that appears to have gamed its scores based on th…
You have to trust that the LLM provider isn't copying the questions when Humanities Last Exam runs the test.
Re: Gemini 3 Pro Model Card [pdf]
#345Earlier quoted context omitted.
Wow. They must have had some major breakthrough. Those scores are truly insane. O_O Models have begun to fairly thoroughly saturate "knowledge" and such, but there are still considerable bumps there But the _big news_, and the demonstration of their achievement here, are the incredible scores they've racked up here for what's necessary for agentic AI to become widely deployable. t2-bench. Visual comprehension. Comput…
The problem is that we know in advance what is the benchmark, so Humanity's Last Exam for example, it's way easier to optimize your model when you have seen the questions before.
Re: Gemini 3 Pro Model Card [pdf]
#346It is interesting that the Gemini 3 beats every other model on these benchmarks, mostly by a wide margin, but not on SWE Bench. Sonnet is still king here and all three look to be basically on the same level. Kind of wild to see them hit such a wall when it comes to agentic coding
This might also hint at SWE struggling to capture what “being good at coding” means. Evals are hard.
Re: Gemini 3 Pro Model Card [pdf]
#347It is interesting that the Gemini 3 beats every other model on these benchmarks, mostly by a wide margin, but not on SWE Bench. Sonnet is still king here and all three look to be basically on the same level. Kind of wild to see them hit such a wall when it comes to agentic coding
From my personal experience using the CLI agentic coding tools, I think gemini-cli is fairly on par with the rest in terms of the planning/code that is generated. However, when I recently tried qwen-code, it gave me a better sense of reasoning and structure that geimini. Claude definitely has it's own advantages but is expensive(at least for some if not for all). My point is, although the model itself may have perfor…
Re: Gemini 3 Pro Model Card [pdf]
#348Earlier quoted context omitted.
I think Anthropic is reading the room, and just going to go hard on being "the" coding model. I suppose they feel that if they can win that, they can get an ROI without having to do full blown multimodality at the highest level. It's probably pretty liberating, because you can make a "spikey" intelligence with only one spike to really focus on.
Codex has been good enough to me and it’s much cheaper. I code non-trivial stuff with it like multi-threaded code and at least for my style of AI coding which is to do fairly small units of work with multiple revisions it is good enough for me to not to even consider the competition. Just giving you a perspective on how the benchmarks might not be important at all for some people and how Claude may have a difficult t…
In Claude on the other hand, MCP connections really do seem to ‘just work’
Re: Gemini 3 Pro Model Card [pdf]
#349Earlier quoted context omitted.
Can you explain what you mean by this? iPhone was the end of Blackberry. It seems reasonable that a smarter, cheaper, faster model would obsolete anything else. ChatGPT has some brand inertia, but not that much given it's barely 2 years old.
Ask yourself why Microsoft Teams won. These are business tools first and foremost.
Re: Gemini 3 Pro Model Card [pdf]
#350Benchmarks from page 4 of the model card: | Benchmark | 3 Pro | 2.5 Pro | Sonnet 4.5 | GPT-5.1 | |-----------------------|-----------|---------|------------|-----------| | Humanity's Last Exam | 37.5% | 21.6% | 13.7% | 26.5% | | ARC-AGI-2 | 31.1% | 4.9% | 13.6% | 17.6% | | GPQA Diamond | 91.9% | 86.4% | 83.4% | 88.1% | | AIME 2025 | | | | | | (no tools) | 95.0% | 88.0% | 87.0% | 94.0% | | (code execution) | 100% | -…
These numbers are impressive, at least to say. It looks like Google has produced a beast that will raise the bar even higher. What's even more impressive is how Google came into this game late and went from producing a few flops to being the leader at this (actually, they already achieved the title with 2.5 Pro). What makes me even more curious is the following > Model dependencies: This model is not a modification o…
The biggest strides in the last 6-8 months have been in generative AIs, specifically for animation.