Live data from Hacker News

Gemini 3 Deep Think

blog.google

61–70 of 722 posts

Re: Gemini 3 Deep Think

#61
post #19
post #15

Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.

Trick? Lol not a chance. Alphabet is a pure play tech firm that has to produce products to make the tech accessible. They really lack in the latter and this is visible when you see the interactions of their VP's. Luckily for them, if you start to create enough of a lead with the tech, you get many chances to sort out the product stuff.

You sound like Russ Hanneman from SV

Re: Gemini 3 Deep Think

#63
post #39

The pelican riding a bicycle is excellent . I think it's the best I've seen. https://simonwillison.net/2026/Feb/12/gemini-3-deep-think/

Do you have to still keep trying to bang on about this relentlessly? It was sort of humorous for the maybe first 2 iterations, now it's tacky, cheesy, and just relentless self-promotion. Again, like I said before, it's also a terrible benchmark.

Eh, i find it more of a not very informative but lighthearted commentary

Re: Gemini 3 Deep Think

#64

Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...

Arc-AGI (and Arc-AGI-2) is the most overhyped benchmark around though. It's completely misnamed. It should be called useless visual puzzle benchmark 2. It's a visual puzzle, making it way easier for humans than for models trained on text firstly. Secondly, it's not really that obvious or easy for humans to solve themselves! So the idea that if an AI can solve "Arc-AGI" or "Arc-AGI-2" it's super smart or even "AGI" is…

The puzzles are calibrated for human solve rates, but otherwise I agree.

Re: Gemini 3 Deep Think

#65
post #26

Earlier quoted context omitted.

Weren't we barely scraping 1-10% on this with state of the art models a year ago and it was considered that this is the final boss, ie solve this and its almost AGI-like? I ask because I cannot distinguish all the benchmarks by heart.

Yes, but benchmarks like this are often flawed because leading model labs frequently participate in 'benchmarkmaxxing' - ie improvements on ARC-AGI2 don't necessarily indicate similar improvements in other areas (though it does seem like this is a step function increase in intelligence for the Gemini line of models)

Could it also be that the models are just a lot better than a year ago?

Re: Gemini 3 Deep Think

#66

According to benchmarks in the announcement, healthily ahead of Claude 4.6. I guess they didn't test ChatGPT 5.3 though. Google has definitely been pulling ahead in AI over the last few months. I've been using Gemini and finding it's better than the other models (especially for biology where it doesn't refuse to answer harmless questions).

Google is way ahead in visual AI and world modelling. They're lagging hard in agentic AI and autonomous behavior.

Re: Gemini 3 Deep Think

#67
post #13
post #10

Earlier quoted context omitted.

Interestingly, the title of that PDF calls it "Gemini 3.1 Pro". Guess that's dropping soon.

I looked at the file name but not the document title (specifically because I was wondering if this is 3.1). Good spot. edit: they just removed the reference to "3.1" from the pdf

I think this is 3.1 (3.0 Pro with the RL improv of 3.0 Flash). But they probably decided to market it as Deep Think because why not charge more for it.

Re: Gemini 3 Deep Think

#68
post #26

Earlier quoted context omitted.

Weren't we barely scraping 1-10% on this with state of the art models a year ago and it was considered that this is the final boss, ie solve this and its almost AGI-like? I ask because I cannot distinguish all the benchmarks by heart.

Yes, but benchmarks like this are often flawed because leading model labs frequently participate in 'benchmarkmaxxing' - ie improvements on ARC-AGI2 don't necessarily indicate similar improvements in other areas (though it does seem like this is a step function increase in intelligence for the Gemini line of models)

Isn’t the point of ARC that you can’t train against it? Or doesn’t it achieve that goal anymore somehow?

Re: Gemini 3 Deep Think

#69
post #40
post #39

The pelican riding a bicycle is excellent . I think it's the best I've seen. https://simonwillison.net/2026/Feb/12/gemini-3-deep-think/

It's worth noting that you mean excellent in terms of prior AI output. I'm pretty sure this wouldn't be considered excellent from a "human made art" perspective. In other words, it's still got a ways to go! Edit: someone needs to explain why this comment is getting downvoted, because I don't understand. Did someone's ego get hurt, or what?

It depends, if you meant from a human coding an SVG "manually" the same way, I'd still say this is excellent (minus the reflection issue). If you meant a human using a proper vector editor, then yeah.

Re: Gemini 3 Deep Think

#70
I can't shake of the feeling that Googles Deep Think Models are not really different models but just the old ones being run with higher number of parallel subagents, something you can do by yourself with their base model and opencode.
Post reply on HN