Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.
Trick? Lol not a chance. Alphabet is a pure play tech firm that has to produce products to make the tech accessible. They really lack in the latter and this is visible when you see the interactions of their VP's. Luckily for them, if you start to create enough of a lead with the tech, you get many chances to sort out the product stuff.
Gemini 3 Deep Think
61–70 of 722 posts
Re: Gemini 3 Deep Think
#62Re: Gemini 3 Deep Think
#63The pelican riding a bicycle is excellent . I think it's the best I've seen. https://simonwillison.net/2026/Feb/12/gemini-3-deep-think/
Do you have to still keep trying to bang on about this relentlessly? It was sort of humorous for the maybe first 2 iterations, now it's tacky, cheesy, and just relentless self-promotion. Again, like I said before, it's also a terrible benchmark.
Re: Gemini 3 Deep Think
#64Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...
Arc-AGI (and Arc-AGI-2) is the most overhyped benchmark around though. It's completely misnamed. It should be called useless visual puzzle benchmark 2. It's a visual puzzle, making it way easier for humans than for models trained on text firstly. Secondly, it's not really that obvious or easy for humans to solve themselves! So the idea that if an AI can solve "Arc-AGI" or "Arc-AGI-2" it's super smart or even "AGI" is…
Re: Gemini 3 Deep Think
#65Earlier quoted context omitted.
Weren't we barely scraping 1-10% on this with state of the art models a year ago and it was considered that this is the final boss, ie solve this and its almost AGI-like? I ask because I cannot distinguish all the benchmarks by heart.
Yes, but benchmarks like this are often flawed because leading model labs frequently participate in 'benchmarkmaxxing' - ie improvements on ARC-AGI2 don't necessarily indicate similar improvements in other areas (though it does seem like this is a step function increase in intelligence for the Gemini line of models)
Re: Gemini 3 Deep Think
#66According to benchmarks in the announcement, healthily ahead of Claude 4.6. I guess they didn't test ChatGPT 5.3 though. Google has definitely been pulling ahead in AI over the last few months. I've been using Gemini and finding it's better than the other models (especially for biology where it doesn't refuse to answer harmless questions).
Re: Gemini 3 Deep Think
#67Earlier quoted context omitted.
Interestingly, the title of that PDF calls it "Gemini 3.1 Pro". Guess that's dropping soon.
I looked at the file name but not the document title (specifically because I was wondering if this is 3.1). Good spot. edit: they just removed the reference to "3.1" from the pdf
Re: Gemini 3 Deep Think
#68Earlier quoted context omitted.
Weren't we barely scraping 1-10% on this with state of the art models a year ago and it was considered that this is the final boss, ie solve this and its almost AGI-like? I ask because I cannot distinguish all the benchmarks by heart.
Yes, but benchmarks like this are often flawed because leading model labs frequently participate in 'benchmarkmaxxing' - ie improvements on ARC-AGI2 don't necessarily indicate similar improvements in other areas (though it does seem like this is a step function increase in intelligence for the Gemini line of models)
Re: Gemini 3 Deep Think
#69The pelican riding a bicycle is excellent . I think it's the best I've seen. https://simonwillison.net/2026/Feb/12/gemini-3-deep-think/
It's worth noting that you mean excellent in terms of prior AI output. I'm pretty sure this wouldn't be considered excellent from a "human made art" perspective. In other words, it's still got a ways to go! Edit: someone needs to explain why this comment is getting downvoted, because I don't understand. Did someone's ego get hurt, or what?