Live data from Hacker News

Gemini 3

blog.google

71–80 of 1001 posts

Re: Gemini 3

#71

Pelican riding a bicycle: https://pasteboard.co/CjJ7Xxftljzp.png

Some time I think I should spend $50 on Upwork to get a real human artist to do it first to know what is that we're going for. What a good pelican riding a bicycle SVG is actually looking like?

Re: Gemini 3

#72

I'm sure this is a very impressive model, but gemini-3-pro-preview is failing spectacularly at my fairly basic python benchmark. In fact, gemini-2.5-pro gets a lot closer (but is still wrong). For reference: gpt-5.1-thinking passes, gpt-5.1-instant fails, gpt-5-thinking fails, gpt-5-instant fails, sonnet-4.5 passes, opus-4.1 passes (lesser claude models fail). This is a reminder that benchmarks are meaningless – you…

and models are still pretty bad at playing tic-tac-toe, they can do it, but think way too much

it's easy to focus on what they can't do

Re: Gemini 3

#74
post #29

Supposedly this is the model card. Very impressive results. https://pbs.twimg.com/media/G6CFG6jXAAA1p0I?format=jpg&name=... Also, the full document: https://archive.org/details/gemini-3-pro-model-card/page/n3/...

Every time I see a table like this numbers go up. Can someone explain what this actually means? Is there just an improvement that some tests are solved in a better way or is this a breakthrough and this model can do something that all others can not?

I estimate another 7 months before models start getting 115% on Humanity's Last Exam.

Re: Gemini 3

#75
DeepMind page: https://deepmind.google/models/gemini/

Gemini 3 Pro DeepMind Page: https://deepmind.google/models/gemini/pro/

Developer blog: https://blog.google/technology/developers/gemini-3-developer...

Gemini 3 Docs: https://ai.google.dev/gemini-api/docs/gemini-3

Google Antigravity: https://antigravity.google/

Re: Gemini 3

#76

Earlier quoted context omitted.

I don't think it would be a good idea to publish it on a prime source of training data.

He could post an encrypted version and post the key with it to avoid it being trained on?

What makes you think it wouldn't end up in the training set anyway?

Re: Gemini 3

#78
post #65

My favorite benchmark is to analyze a very long audio file recording of a management meeting and produce very good notes along with a transcript labeling all the speakers. 2.5 was decently good at generating the summary, but it was terrible at labeling speakers. 3.0 has so far absolutely nailed speaker labeling.

What prompt do you use for that?

Re: Gemini 3

#79

I'm sure this is a very impressive model, but gemini-3-pro-preview is failing spectacularly at my fairly basic python benchmark. In fact, gemini-2.5-pro gets a lot closer (but is still wrong). For reference: gpt-5.1-thinking passes, gpt-5.1-instant fails, gpt-5-thinking fails, gpt-5-instant fails, sonnet-4.5 passes, opus-4.1 passes (lesser claude models fail). This is a reminder that benchmarks are meaningless – you…

and models are still pretty bad at playing tic-tac-toe, they can do it, but think way too much it's easy to focus on what they can't do

[deleted]

Re: Gemini 3

#80
Curious to see it in action. Gemini 2.5 has already been very impressive as a study buddy for courses like set theory, information theory, and automata. Although I’m always a bit skeptical of these benchmarks. Seems quite unlikely that all of the questions remain out of their training data.
Post reply on HN