Pelican riding a bicycle: https://pasteboard.co/CjJ7Xxftljzp.png
Gemini 3
71–80 of 1001 posts
Re: Gemini 3
#72I'm sure this is a very impressive model, but gemini-3-pro-preview is failing spectacularly at my fairly basic python benchmark. In fact, gemini-2.5-pro gets a lot closer (but is still wrong). For reference: gpt-5.1-thinking passes, gpt-5.1-instant fails, gpt-5-thinking fails, gpt-5-instant fails, sonnet-4.5 passes, opus-4.1 passes (lesser claude models fail). This is a reminder that benchmarks are meaningless – you…
it's easy to focus on what they can't do
Re: Gemini 3
#73Re: Gemini 3
#74Supposedly this is the model card. Very impressive results. https://pbs.twimg.com/media/G6CFG6jXAAA1p0I?format=jpg&name=... Also, the full document: https://archive.org/details/gemini-3-pro-model-card/page/n3/...
Every time I see a table like this numbers go up. Can someone explain what this actually means? Is there just an improvement that some tests are solved in a better way or is this a breakthrough and this model can do something that all others can not?
Re: Gemini 3
#75Gemini 3 Pro DeepMind Page: https://deepmind.google/models/gemini/pro/
Developer blog: https://blog.google/technology/developers/gemini-3-developer...
Gemini 3 Docs: https://ai.google.dev/gemini-api/docs/gemini-3
Google Antigravity: https://antigravity.google/
Re: Gemini 3
#76Re: Gemini 3
#77Pelican riding a bicycle: https://pasteboard.co/CjJ7Xxftljzp.png
Re: Gemini 3
#78My favorite benchmark is to analyze a very long audio file recording of a management meeting and produce very good notes along with a transcript labeling all the speakers. 2.5 was decently good at generating the summary, but it was terrible at labeling speakers. 3.0 has so far absolutely nailed speaker labeling.
Re: Gemini 3
#79I'm sure this is a very impressive model, but gemini-3-pro-preview is failing spectacularly at my fairly basic python benchmark. In fact, gemini-2.5-pro gets a lot closer (but is still wrong). For reference: gpt-5.1-thinking passes, gpt-5.1-instant fails, gpt-5-thinking fails, gpt-5-instant fails, sonnet-4.5 passes, opus-4.1 passes (lesser claude models fail). This is a reminder that benchmarks are meaningless – you…
and models are still pretty bad at playing tic-tac-toe, they can do it, but think way too much it's easy to focus on what they can't do