Gemini 3 Pro Model Card [pdf]
301–310 of 359 posts
Re: Gemini 3 Pro Model Card [pdf]
#302https://storage.googleapis.com/deepmind-media/Model-Cards/Ge...
Re: Gemini 3 Pro Model Card [pdf]
#303Earlier quoted context omitted.
Used an AI to populate some of 5.1 thinking's results. Benchmark | Gemini 3 Pro | Gemini 2.5 Pro | Claude Sonnet 4.5 | GPT-5.1 | GPT-5.1 Thinking ---------------------------|--------------|----------------|-------------------|---------|------------------ Humanity's Last Exam | 37.5% | 21.6% | 13.7% | 26.5% | 52% ARC-AGI-2 | 31.1% | 4.9% | 13.6% | 17.6% | 28% GPQA Diamond | 91.9% | 86.4% | 83.4% | 88.1% | 61% AIM 2025…
What? The 4.5 and 5.1 columns aren't thinking in Google's report? That's a scandal, IMO. Given that Gemini-3 seems to do "fine" against the thinking versions why didn't they post those results? I get that PMs like to make a splash but that's shockingly dishonest.
> For Claude Sonnet 4.5, and GPT-5.1 we default to reporting high reasoning results, but when reported results are not available we use best available reasoning results.
https://storage.googleapis.com/deepmind-media/gemini/gemini_...
Re: Gemini 3 Pro Model Card [pdf]
#304Earlier quoted context omitted.
The problem is that we know in advance what is the benchmark, so Humanity's Last Exam for example, it's way easier to optimize your model when you have seen the questions before.
From https://lastexam.ai/ : "The dataset consists of 2,500 challenging questions across over a hundred subjects. We publicly release these questions, while maintaining a private test set of held out questions to assess model overfitting ." [emphasis mine] While the private questions don't seem to be included in the performance results, HLE will presumably flag any LLM that appears to have gamed its scores based on th…
This was the primary bottleneck preventing models from tackling novel scientific problems they haven't seen before.
If Gemini 3 Pro has transcended "reading the internet" (knowledge saturation), and made huge progress in "thinking about the internet" (reasoning scaling), then this is a really big deal.
Re: Gemini 3 Pro Model Card [pdf]
#305Earlier quoted context omitted.
Claude 3.5 came out in June of last year, and it is imo marginally worse than the AI models currently available for coding. I do not think models are 10x better than 1 year ago, that seems extremely hyperbolic or you are working in a super niche area where that is true.
Are you using it for agentic tasks of any length? 3.5 and 4.5 are about the same for single file/single snippet tasks, but my observation has been that 4.5 can do longer, more complex tasks that were a waste of time to even try with 3.5 because it would always fail.
Re: Gemini 3 Pro Model Card [pdf]
#306Earlier quoted context omitted.
The problem is that we know in advance what is the benchmark, so Humanity's Last Exam for example, it's way easier to optimize your model when you have seen the questions before.
From https://lastexam.ai/ : "The dataset consists of 2,500 challenging questions across over a hundred subjects. We publicly release these questions, while maintaining a private test set of held out questions to assess model overfitting ." [emphasis mine] While the private questions don't seem to be included in the performance results, HLE will presumably flag any LLM that appears to have gamed its scores based on th…
Re: Gemini 3 Pro Model Card [pdf]
#307Re: Gemini 3 Pro Model Card [pdf]
#308Earlier quoted context omitted.
The problem is that we know in advance what is the benchmark, so Humanity's Last Exam for example, it's way easier to optimize your model when you have seen the questions before.
From https://lastexam.ai/ : "The dataset consists of 2,500 challenging questions across over a hundred subjects. We publicly release these questions, while maintaining a private test set of held out questions to assess model overfitting ." [emphasis mine] While the private questions don't seem to be included in the performance results, HLE will presumably flag any LLM that appears to have gamed its scores based on th…
Re: Gemini 3 Pro Model Card [pdf]
#309Earlier quoted context omitted.
From https://lastexam.ai/ : "The dataset consists of 2,500 challenging questions across over a hundred subjects. We publicly release these questions, while maintaining a private test set of held out questions to assess model overfitting ." [emphasis mine] While the private questions don't seem to be included in the performance results, HLE will presumably flag any LLM that appears to have gamed its scores based on th…
How do they hold back questions in practice though? These are hosted models. To ask the question is to reveal it to the model team.
Re: Gemini 3 Pro Model Card [pdf]
#310Earlier quoted context omitted.
The problem is that we know in advance what is the benchmark, so Humanity's Last Exam for example, it's way easier to optimize your model when you have seen the questions before.
From https://lastexam.ai/ : "The dataset consists of 2,500 challenging questions across over a hundred subjects. We publicly release these questions, while maintaining a private test set of held out questions to assess model overfitting ." [emphasis mine] While the private questions don't seem to be included in the performance results, HLE will presumably flag any LLM that appears to have gamed its scores based on th…