Live data from Hacker News

Gemini 3.1 Pro

blog.google

701–710 of 951 posts

Re: Gemini 3.1 Pro

#701

People underrate Google's cost effectiveness so much. Half price of Opus. HALF. Think about ANY other product and what you'd expect from the competition thats half the price. Yet people here act like Gemini is dead weight ____ Update: 3.1 was 40% of the cost to run AA index vs Opus Thinking AND SONNET, beat Opus, and still 30% faster for output speed. https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...

You can pay 1 cent for a mediocre answer or 2 cents for a great answer. So a lot of these things are relative. Now if that equation plays out 20K times a day, well that's one thing, but if it's 'once a day' then the cost basis becomes irrelevant. Like the cost of staplers for the Medical Device company. Obviously it will matter, but for development ... it's probably worth it to pay $300/mo for the best model, when th…

Right now I'll pay 2x for a subjectively 20+% better coding agent. But in a year I don't think there will be an agent that to me is subjectively 20% better amongst the big three.

Re: Gemini 3.1 Pro

#702

Earlier quoted context omitted.

Google are stuck because they have to compete with OpenAI. If they don’t, they face an existential threat to their advertising business. But then they leave the door open for Anthropic on coding, enterprise and agentic workflows. Sensibly, that’s what they seem to be doing. That said Gemini is noticeably worse than ChatGPT (it’s quite erratic) and Anthropic’s work on coding / reasoning seems to be filtering back to i…

Google might be a mess now, but they have time. OpenAI and Anthropic are on barrowed time, Google has a built in money printer. They just need to outlast the others.

Plus they started making AI processors 11 years ago and invented the math behind “GPTs” 9 years ago. Gemini is way cheaper to run for them than it does for everyone else.

I think Gemini is really built for their biggest market — Google Search. You ask questions and get answers.

I’m sure they’ll figure out agentic flows. Google is always a mess when it comes to product. Don’t forget the Google chat sagas where it seems as if different parts of the company were making the same product.

Re: Gemini 3.1 Pro

#703

I hope this works better than 3.0 Pro I'm a former Googler and know some people near the team, so I mildly root for them to at least do well, but Gemini is consistently the most frustrating model I've used for development. It's stunningly good at reasoning, design, and generating the raw code, but it just falls over a lot when actually trying to get things done, especially compared to Claude Opus. Within VS Code Copi…

Yep, great models to use in gemini.google.com but outside of that it somehow becomes dumb (especially for coding)

Re: Gemini 3.1 Pro

#704

Earlier quoted context omitted.

Spark is the 'same model and harness' but on Cerebras. Your intuition may be deceiving you, maybe assuming it's a speed/quality trade-off, it's not. It's just faster hardware. No IQ tradeoff. If you toy around with Cerebras directly, you get a feel for it. Edit: see note below, I'm wrong. Not same model.

> Today, we’re releasing a research preview of GPT‑5.3‑Codex‑Spark, a smaller version of GPT‑5.3‑Codex , and our first model designed for real-time coding. from https://openai.com/index/introducing-gpt-5-3-codex-spark/ , emphasis mine

You're right. It's funny because I kind of noticed that, but with all of these subtle model issues, I'm so used to being distraught by the smallest thing I've had to learn to 'trust the data' aka the charts, model standings, performance, etc. and in this case, I was under the assumption 'it was the same model' clearly it's not.

Which is a bummer because it would be nice to try a true side-by-side analysis.

Re: Gemini 3.1 Pro

#705

Earlier quoted context omitted.

Yes, this is very true and it speaks strongly to this wayward notion of 'models' - it depends so much on the tuning, the harness, the tools. I think it speaks to the broader notion of AGI as well. Claude is definitively trained on the process of coding not just the code, that much is clear. Codex has the same limitation but not quite as bad. This may be a result of Anthropic using 'user cues' with respect to what are…

Google are stuck because they have to compete with OpenAI. If they don’t, they face an existential threat to their advertising business. But then they leave the door open for Anthropic on coding, enterprise and agentic workflows. Sensibly, that’s what they seem to be doing. That said Gemini is noticeably worse than ChatGPT (it’s quite erratic) and Anthropic’s work on coding / reasoning seems to be filtering back to i…

They all suck!!!

Re: Gemini 3.1 Pro

#706

People underrate Google's cost effectiveness so much. Half price of Opus. HALF. Think about ANY other product and what you'd expect from the competition thats half the price. Yet people here act like Gemini is dead weight ____ Update: 3.1 was 40% of the cost to run AA index vs Opus Thinking AND SONNET, beat Opus, and still 30% faster for output speed. https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...

You can pay 1 cent for a mediocre answer or 2 cents for a great answer. So a lot of these things are relative. Now if that equation plays out 20K times a day, well that's one thing, but if it's 'once a day' then the cost basis becomes irrelevant. Like the cost of staplers for the Medical Device company. Obviously it will matter, but for development ... it's probably worth it to pay $300/mo for the best model, when th…

Yeah you’re right but most people in the world do not need an agent that codes.

I think Gemini gives fine answers outside code tasks.

Outside of work, where I use Claude, Gemini is cheaper for me (for what I would use AI for) than both Claude and ChatGPT so Google gets my money.

Re: Gemini 3.1 Pro

#707
I think we're past the point where benchmarks hold real value. All models are above a certain threshold of intelligence but Gemini somehow borrows the worst of both worlds. It's neither good with long-horizon coding tasks nor does it offer a likable personality (like Claude which is much more beloved)

Re: Gemini 3.1 Pro

#708

Earlier quoted context omitted.

We are not at the moment where price matters. All that matters is performance.

What did you say? Cant hear you over the $400B in capex spend. Counterpoint: price will matter before we hit AGI

Why do you believe it has to? Uber took 15 years to show a profit. 15 years from 2022 when chatgpt launched is 2037. That's long enough that to say I don't know if I'll even be alive by then.

Re: Gemini 3.1 Pro

#709
What I’m noticing, overall: I’ve never cut so much code in my life. I’ve become a coding monster with one of those dark green GitHub profiles ever since 5.3-Codex gave me the confidence to load in a ridiculous number of tasks every day and let it rip. I have about three coding tasks going at once and in another window, Claude Cowork is ripping through PowerPoints and getting back to lawyers.

This tech is not going to replace us. If anything, I am becoming even more of a workaholic. But the output volume is going to pay off for those who are privileged enough to use these tools.

Re: Gemini 3.1 Pro

#710
post #709

What I’m noticing, overall: I’ve never cut so much code in my life. I’ve become a coding monster with one of those dark green GitHub profiles ever since 5.3-Codex gave me the confidence to load in a ridiculous number of tasks every day and let it rip. I have about three coding tasks going at once and in another window, Claude Cowork is ripping through PowerPoints and getting back to lawyers. This tech is not going to…

Yeah see this article I think it was spot on

https://hbr.org/2026/02/ai-doesnt-reduce-work-it-intensifies...

Post reply on HN