Live data from Hacker News

Gemini 3.1 Pro

blog.google

781–790 of 951 posts

Re: Gemini 3.1 Pro

#781
post #772

Earlier quoted context omitted.

Yup, you got it. It's a weird situation for sure. You know what's also weird: Gem3 'Pro' is pretty dumb. OAI has 'thinking levels' which work pretty well, it's nice to have the 'super duper' button - but also - they have the 'Pro' product which is another model altogether and thinks for 20 min. It's different than 'Research'. OAI Pro (+ maybe Spark) is the only reason I have OAI sub. Neither Anthropic nor Google seem…

Can you explain what’s so different about pro? I’ve used everything frontier model and had Pro a while ago but it seemed to just be the same models served faster at the time.

It's a different model and designed to 'think very hard' about issues. It's basically a 'very extended thinking mixed with research' type of solution.

While the 'research' solutions tend to go very wide and come back with a 'paper' the Pro model seems to do an exhaustive amount of thinking combined with research, and tries to integrate findings. I think it goes down a lot of rabbit holes.

I find it's by far the best way to find solutions to hard problems, but it typically does require a 'hard problem' in order to shine.

And it takes an enormous amount of time. Ito could be essentially a form of 'saturating the problem with tokens'. It's OAI's most expensive model by far. A prompt usually costs me $1-3 if paying per token.

Re: Gemini 3.1 Pro

#782
Gemini is the smartest model currently available. It is the only model out of the big ones that correcly identifies the specific versions of superhers in a collage I tested them with.

Re: Gemini 3.1 Pro

#783

Earlier quoted context omitted.

Plus they started making AI processors 11 years ago and invented the math behind “GPTs” 9 years ago. Gemini is way cheaper to run for them than it does for everyone else. I think Gemini is really built for their biggest market — Google Search. You ask questions and get answers. I’m sure they’ll figure out agentic flows. Google is always a mess when it comes to product. Don’t forget the Google chat sagas where it seem…

Why do you assume they’ll figure it out when they pretty consistently mess things up?

How do they consistently mess things up ? Current market cap 3.7T, only Apple and Nvidia are bigger. Youtube is a huge success, Search is still growing at 10%-15% which is crazy, cloud growing at 35%ish, TPUs enable them to be independent from NVidia etc. Gemini market share went up from 5%-6% early 2025 to 21% early 2026. I personally bet Gemini market share will keep growing. They are executing well on all verticals imo, not messing up.

Re: Gemini 3.1 Pro

#784
post #758

Earlier quoted context omitted.

I mean their ads business just broke $80b per quarter, not sure where this idea is coming from...

Google hasn't seen its legacy ad revenue start to dent until products with built-in agents start to see mass adoption. Writing is on the wall that orders of magnitude fewer people will be going to google.com or using an interactive Google search in the next 5 years though.

LLMs are pretty mediocre for a lot of money queries like searching to buy shoes, looking at flights etc due to them not being up to date. So sure you can use them as a wrapper on top of Google but I assume a huge chunk of people will just go to Google to do that or use Google agents. Chrome will prove a very valuable asset for that - the whole experience can become agentic and Google is very well positioend to convert billions of users into their AI. Power of habit and also Google will deliver a very high quality experience at scale that only OpenAI can currently compete with. I'm not saying their search / ads revenue is never gonna drop - it might. But it will be a slow process (as we can see. it's actually still freaking growing in the high tens) and Google is well positioned to recover the lost revenue with its A.I offerings.

Re: Gemini 3.1 Pro

#785

Earlier quoted context omitted.

I was being facetious, I mean one day models might skip the middle man of code and compilation and take your specs and produce an ultra efficent binary.

Musk was saying that recently but I don't see it being efficient or worthwhile to do this. I could be proven brutally wrong, but code is language; executables aren't. There's also no real reason to bother with this when we have quick-compiling languages. More realistically, I could see particular languages and frameworks proving out to be more well-designed and apt for AI code creation; for instance, I was always too…

I’ve thought an interesting outcome might be that it’s not even that there’s a binary generated. It’s just user input -> machine code LLM -> CPU. Like the only binary would be the LLM itself and it’s essentially mimicking software live. The paper “Diffusion as a Model of Environment Dream” (DIAMOND) is close to what I’m thinking, where they have a diffusion model generate frames of a game, updating with user input, but there’s no actual “game” code it’s just the model.

https://diamond-wm.github.io/

Like you’d have a machine code LLM that behaves like software but instead of a static binary being executed it’s just the LLM itself “executing” on inputs and precious state. I’m horrible at communicating this idea but hopefully the gist is there.

Re: Gemini 3.1 Pro

#787
post #486

Earlier quoted context omitted.

ARC 2 was made specifically to artificially lower contemporary LLM scores, therefore any kind of model improvements will have outsized effects Also people use "saturated" too liberally. The top left corner 1 cent per task is saturated IMO. Since there are billions of people who would perfer to solve arc 1 tasks at 52 cents per task. Arc 2 a human would make thousands of dollars a day with 99.99% accuracy

You are saying something interesting but too esoteric. Can you explain for beginners?

You could get rich by solving ARC 2 tasks yourself instead of forwarding the work to an LLM, given a client willing to pay LLM rate.

Re: Gemini 3.1 Pro

#788

You know what would slay right now? A native app. Not another piece of Electron bloatware, a regular, efficient, fast, snappy, native, app. One that connects to my MCP severs and has local filesystem tools. Anthropic might fall behind Google/OpenAI eventually, but their Desktop App + MCP/Connectors is unbelievably useful to get real work done.

I haven't used Anthropic's desktop app in months since I don't have access to a Mac anymore, but when I did...it was just an electron app? Did something change?

Not only that, it is the slowest app among all AI apps.

Re: Gemini 3.1 Pro

#789

People underrate Google's cost effectiveness so much. Half price of Opus. HALF. Think about ANY other product and what you'd expect from the competition thats half the price. Yet people here act like Gemini is dead weight ____ Update: 3.1 was 40% of the cost to run AA index vs Opus Thinking AND SONNET, beat Opus, and still 30% faster for output speed. https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...

Gemini is the most paradoxical model because it benchmarks great even in private benchmarks done by regular people, Deep Mind is unquestionably full of capable engineers with incredible skill, and personally Gemini has been great for my day job and my coding for fun (not for profit) endeavors. Switching between it and 4.6 in antigravity and I don't see much of a difference, they both do what I ask. But man, people ar…

I personally found Gemini 3.0 to step on my toes in Agentic coding. I tried it around 10 or so times but it quickly became apparent that it was somehow coming to its own conclusions about what needs to be done instead of following instructions.

Like files I didn't mention being edited and read and stuff of that nature. Sometimes this is cute in fixing typos in docs but when its changing things where it clearly doesn't even understand the intentionality behind something it's annoying.

Gemini 3.1 is clearly much better when trying it today. It stayed focused and found its way around without getting distracted.

Re: Gemini 3.1 Pro

#790

It got the car wash question perfectly: You are definitely going to have to drive it there—unless you want to put it in neutral and push! While 200 feet is a very short and easy walk, if you walk over there without your car, you won't have anything to wash once you arrive. The car needs to make the trip with you so it can get the soap and water. Since it's basically right next door, it'll be the shortest drive of you…

And Gemini 3 can’t..? Isn’t this just a thinking vs nonthinking model thing?
Post reply on HN