If these numbers are true then OpenAI is probably done, Anthropic too. Still, it's hard to see an effective monetization method for this tech and it clearly is eating Google's main pie which is search.
For SWE it is the same ranking. But if Google's $20/mo plan is comparable to the $100-200 plans for OpenAI and Anthropic, yes they are done. But we'll have to wait a few weeks to see if the nerfed model post-release is still as good.
Gemini 3 Pro Model Card [pdf]
161–170 of 359 posts
Re: Gemini 3 Pro Model Card [pdf]
#162Earlier quoted context omitted.
Noteworthily, although Gemini 3 Pro seems to have much benchmark scores than other models across the board (including compared to Claude), it's not the case for coding, where it appears to score essentially the same as the others. I wonder why that is. So far, IMHO, Claude Code remains significantly better than Gemini CLI. We'll see whether that changes with Gemini 3.
from my experience, the quality of gemini-cli isn't great, experiencing lot of stupied bug.
Not that Google didn't use to have problems shipping useful things. But it's gotten a lot worse.
Re: Gemini 3 Pro Model Card [pdf]
#163Earlier quoted context omitted.
The reported results where GPT 5.1 beats Gemini 3 are on SWE Bench Verified, and GPT 5.1 Codex also beats Gemini 3 on Terminal Bench.
You're right on SWE Bench Verified, I missed that and I'll delete my comment. GPT 5.1 Codex beats Gemini 3 on Terminal Bench specifically on Codex CLI, but that's apples-to-oranges (hard to tell how much of that is a Codex-specific harness vs model). Look forward to seeing the apples-to-apples numbers soon, but I wouldn't be surprised if Gemini 3 wins given how close it comes in these benchmarks.
Re: Gemini 3 Pro Model Card [pdf]
#164Curiously, this website seems to be blocked in Spain for whatever reason, and the website's certificate is served by `allot.com/emailAddress=info@allot.com` which obviously fails... Anyone happen to know why? Is this website by any change sharing information on safe medical abortions or women's rights, something which has gotten websites blocked here before?
Re: Gemini 3 Pro Model Card [pdf]
#165Earlier quoted context omitted.
Also does not beat GPT-5.1 Codex on terminal bench (57.8% vs 54.2%): https://www.tbench.ai/ I did not bother verifying the other claims.
Not apples-to-apples. "Codex CLI (GPT-5.1-Codex)", which the site refers to, adds a specific agentic harness, whereas the Gemini 3 Pro seems to be on a standard eval harness. It would be interesting to see the apples-to-apples figure, i.e. with Google's best harness alongside Codex CLI.
What do you mean by "standard eval harness"?
Re: Gemini 3 Pro Model Card [pdf]
#166Curious to see the API pricing. SOTA performance across tasks at a price cheaper than GPT 5 / Claude would make mostly everyone switch to Gemini.
Re: Gemini 3 Pro Model Card [pdf]
#167I know this is a little controversial but the lack of performance on SWE-bench is hugely disappointing I think economically. These models don’t have any viable path to profitability if they can’t take engineering jobs.
I thought that but it does do a lot better on other benchmarks. Perhaps SWE bench just doesn't capture a lot of the improvement? If the web design improvements people have been posting on twitter, I suspect this will be a huge boon for developers. SWE benchmark is really testing bugfixing/feature dev more. Anyway let's see. I'm still hyped!
Re: Gemini 3 Pro Model Card [pdf]
#168Earlier quoted context omitted.
Agreed, too early to write off others entirely. It'll be interesting to see who comes out the other side of the bubble with a working business.
Anthropic has a fairly significant lead when it comes to enterprise usage and for coding. This seems like a workable business model to me.
Re: Gemini 3 Pro Model Card [pdf]
#169It is interesting that the Gemini 3 beats every other model on these benchmarks, mostly by a wide margin, but not on SWE Bench. Sonnet is still king here and all three look to be basically on the same level. Kind of wild to see them hit such a wall when it comes to agentic coding
Re: Gemini 3 Pro Model Card [pdf]
#170It's hilarious that the release of Gemini 3 is getting eclipsed by this cloudflare outage.