I wonder how significant this is. DeepMind was always more research-oriented that OpenAI, which mostly scaled things up. They may have come up with a significantly better architecture (Transformer MoE still leaves a lot of room).
Gemini 3 Pro Model Card [pdf]
121–130 of 359 posts
Re: Gemini 3 Pro Model Card [pdf]
#122Earlier quoted context omitted.
[comment removed]
The reported results where GPT 5.1 beats Gemini 3 are on SWE Bench Verified, and GPT 5.1 Codex also beats Gemini 3 on Terminal Bench.
GPT 5.1 Codex beats Gemini 3 on Terminal Bench specifically on Codex CLI, but that's apples-to-oranges (hard to tell how much of that is a Codex-specific harness vs model). Look forward to seeing the apples-to-apples numbers soon, but I wouldn't be surprised if Gemini 3 wins given how close it comes in these benchmarks.
Re: Gemini 3 Pro Model Card [pdf]
#123Earlier quoted context omitted.
Also does not beat GPT-5.1 Codex on terminal bench (57.8% vs 54.2%): https://www.tbench.ai/ I did not bother verifying the other claims.
Not apples-to-apples. "Codex CLI (GPT-5.1-Codex)", which the site refers to, adds a specific agentic harness, whereas the Gemini 3 Pro seems to be on a standard eval harness. It would be interesting to see the apples-to-apples figure, i.e. with Google's best harness alongside Codex CLI.
Re: Gemini 3 Pro Model Card [pdf]
#124Re: Gemini 3 Pro Model Card [pdf]
#125Earlier quoted context omitted.
At least at the moment, coming in late seems to matter little. Anyone with money can trivially catch up to a state of the art model from six months ago. And as others have said, late is really a function of spigot, guardrails, branding, and ux, as much as it is being a laggard under the hood.
> Anyone with money can trivially catch up to a state of the art model from six months ago. How come apple is struggling then?
The may want to use 3rd party or just wait for AI to be more stable to see how people actually use it instead of adding slop in the core of their product.
Re: Gemini 3 Pro Model Card [pdf]
#126Benchmarks from page 4 of the model card: | Benchmark | 3 Pro | 2.5 Pro | Sonnet 4.5 | GPT-5.1 | |-----------------------|-----------|---------|------------|-----------| | Humanity's Last Exam | 37.5% | 21.6% | 13.7% | 26.5% | | ARC-AGI-2 | 31.1% | 4.9% | 13.6% | 17.6% | | GPQA Diamond | 91.9% | 86.4% | 83.4% | 88.1% | | AIME 2025 | | | | | | (no tools) | 95.0% | 88.0% | 87.0% | 94.0% | | (code execution) | 100% | -…
This is a big jump in most benchmarks.And if it can match other models in coding while having that Google TPM inference speed and the actually native 1m context window, it's going to be a big hit. I hope it's isn't such a sycophant like the current gemini 2.5 models, it makes me doubt its output, which is maybe a good thing now that I think about it.
Its not over and never will be for 2 decade old accounting software, it is definitely will not be over for other AI labs.
Re: Gemini 3 Pro Model Card [pdf]
#127Earlier quoted context omitted.
This is a big jump in most benchmarks.And if it can match other models in coding while having that Google TPM inference speed and the actually native 1m context window, it's going to be a big hit. I hope it's isn't such a sycophant like the current gemini 2.5 models, it makes me doubt its output, which is maybe a good thing now that I think about it.
> it's over for the other labs. What's with the hyperbole? It'll tighten the screws, but saying that it's "over for the other labs' might be a tad premature.
Re: Gemini 3 Pro Model Card [pdf]
#128It is interesting that the Gemini 3 beats every other model on these benchmarks, mostly by a wide margin, but not on SWE Bench. Sonnet is still king here and all three look to be basically on the same level. Kind of wild to see them hit such a wall when it comes to agentic coding
My point is, although the model itself may have performed in benchmarks, I feel like there are other tools that are doing better just by adapting better training/tooling. Gemini cli, in particular, is not so great looking up for latest info on web. Qwen seemed to be trained better around looking up for information (or to reason when/how to), in comparision. Even the step-wise break down of work felt different and a bit smoother.
I do, however, use gemini cli for the most part just because it has a generous free quota with very few downsides comparted to others. They must be getting loads of training data :D.
Re: Gemini 3 Pro Model Card [pdf]
#129I know this is a little controversial but the lack of performance on SWE-bench is hugely disappointing I think economically. These models don’t have any viable path to profitability if they can’t take engineering jobs.
Re: Gemini 3 Pro Model Card [pdf]
#130I saw this on Reddit earlier today. Over there the source of this file was given as: https://web.archive.org/web/20251118111103/https://storage.g... The bucket name "deepmind-media" has been used in the past on the deepmind official site, so it seems legit.