Earlier quoted context omitted.
That is pretty impressive. So impressive it makes you wonder if someone has noticed it being used a benchmark prompt.
Simon says if he gets a suspiciously good result he'll just try a bunch of other absurd animal/vehicle combinations to see if they trained a special case: https://simonwillison.net/2025/Nov/13/training-for-pelicans-...
Gemini 3
61–70 of 1001 posts
Re: Gemini 3
#62Re: Gemini 3
#63Temperature continues to be gated to maximum of 0.2, and there's still the hidden top_k of 64 that you can't turn off.
I love the google AI studio, but I hate it too for not enabling a whole host of advanced features. So many mixed feelings, so many unanswered questions, so many frustrating UI decisions on a tool that is ostensibly aimed at prosumers...
Re: Gemini 3
#64Re: Gemini 3
#65Re: Gemini 3
#66Earlier quoted context omitted.
That is pretty impressive. So impressive it makes you wonder if someone has noticed it being used a benchmark prompt.
Simon says if he gets a suspiciously good result he'll just try a bunch of other absurd animal/vehicle combinations to see if they trained a special case: https://simonwillison.net/2025/Nov/13/training-for-pelicans-...
Re: Gemini 3
#67Re: Gemini 3
#68Supposedly this is the model card. Very impressive results. https://pbs.twimg.com/media/G6CFG6jXAAA1p0I?format=jpg&name=... Also, the full document: https://archive.org/details/gemini-3-pro-model-card/page/n3/...
The person also claims that with thinking on the gap narrows considerably.
We'll probably have 3rd party benchmarks in a couple of days.
Re: Gemini 3
#69Re: Gemini 3
#70I'm sure this is a very impressive model, but gemini-3-pro-preview is failing spectacularly at my fairly basic python benchmark. In fact, gemini-2.5-pro gets a lot closer (but is still wrong). For reference: gpt-5.1-thinking passes, gpt-5.1-instant fails, gpt-5-thinking fails, gpt-5-instant fails, sonnet-4.5 passes, opus-4.1 passes (lesser claude models fail). This is a reminder that benchmarks are meaningless – you…
I just chatted with gemini-3-pro-preview about an idea I had and I'm glad that I did. I will definitely come back to it.
IMHO, the current batch of free, free-ish models are all perfectly adequate for my uses, which are mostly coding, troubleshooting and learning/research.
This is an amazing time to be alive and the AI bubble doomers that are costing me some gains RN can F-Off!