Live data from Hacker News

Gemini 3.1 Pro

blog.google

851–860 of 951 posts

Re: Gemini 3.1 Pro

#851

Earlier quoted context omitted.

I personally found Gemini 3.0 to step on my toes in Agentic coding. I tried it around 10 or so times but it quickly became apparent that it was somehow coming to its own conclusions about what needs to be done instead of following instructions. Like files I didn't mention being edited and read and stuff of that nature. Sometimes this is cute in fixing typos in docs but when its changing things where it clearly doesn'…

The only cases where I've had gemini step on my toes like that is when a) I realized my instructions were unclear or missing something b) my assumptions/instructions were flawed about how/why something needed to be done.

Instruction following has improved a lot since a few years ago but let's not pretend these things are perfect mate.

There's a certain capacity of instructions, albiet its quite high, at which point you will find them skipping points and drifting. It doesn't have to be ambiguity in instructions.

Re: Gemini 3.1 Pro

#852

Earlier quoted context omitted.

Yes, this is very true and it speaks strongly to this wayward notion of 'models' - it depends so much on the tuning, the harness, the tools. I think it speaks to the broader notion of AGI as well. Claude is definitively trained on the process of coding not just the code, that much is clear. Codex has the same limitation but not quite as bad. This may be a result of Anthropic using 'user cues' with respect to what are…

Google are stuck because they have to compete with OpenAI. If they don’t, they face an existential threat to their advertising business. But then they leave the door open for Anthropic on coding, enterprise and agentic workflows. Sensibly, that’s what they seem to be doing. That said Gemini is noticeably worse than ChatGPT (it’s quite erratic) and Anthropic’s work on coding / reasoning seems to be filtering back to i…

I would agree that Gemini is not keeping up with Anthropic on coding, but I completely disagree on ChatGPT. It's been months for me since I've gotten anything from OpenAI that felt like it was worth my time. I don't really consider them anymore.

Google is mostly doing what they've always done. They've created a few tools like Gemini and NotebookLM, and they're going to focus more effort on whatever gets the most traffic. Then anything they can't monetize will get cut.

Re: Gemini 3.1 Pro

#853
post #751

Earlier quoted context omitted.

In my experience Gemini 3.0 pro is noticeably better than chatgpt 5.2 for non-coding tasks. The latter gives me blatantly wrong information all the time, the former very rarely.

Strange that you say that because the general consensus (and my experience) seems to be the opposite, as well as the AA-Omniscience Hallucination Rate Benchmark which puts 3.0 Pro among the higher hallucinating models. 3.1 seems to be a noticeable improvement though.

> the AA-Omniscience Hallucination Rate Benchmark which puts 3.0 Pro among the higher hallucinating models. 3.1 seems to be a noticeable improvement though.

As sibling comment says, AA-Omniscience Hallucination Rate Benchmark puts Gemini 3.0 as the best performing aside from Gemini 3.1 preview.

https://artificialanalysis.ai/evaluations/omniscience

Re: Gemini 3.1 Pro

#854

Earlier quoted context omitted.

Gemini is the most paradoxical model because it benchmarks great even in private benchmarks done by regular people, Deep Mind is unquestionably full of capable engineers with incredible skill, and personally Gemini has been great for my day job and my coding for fun (not for profit) endeavors. Switching between it and 4.6 in antigravity and I don't see much of a difference, they both do what I ask. But man, people ar…

People can be and often are wrong. You'd notice how good Opus is in Claude Code. IMHO CC is the secret sauce

> IMHO CC is the secret sauce

Cant smart people just reverse engineer CC and figure out what is the secret sauce atleast for CC App?

Re: Gemini 3.1 Pro

#855

Earlier quoted context omitted.

There are thousands like you now. How many does it take to run the economy? What would the rest do. Think of it like what a tractor did to agricultural work. The fist guy that used a tractor probably thought: this is not replacing me, I’m just much more productive. Well, turns out you only need one guy per farm now.

But now many suburban homeowners also have a little lawn tractor, and lots of people on small acreage have a utility tractor. None of them are farmers, but they get value out of the technology as well. Plus, we're feeding a lot more people for a lot less money than we did before tractors.

Yeah, but we used to employ hundreds of people per farm, or per plantation, to be exact. Thousands maybe to do the sugar cane work, as an example. Replaced by 5 high tech, GPS driven, human on board to supervise, not even to drive, tractors.

So human doing lawn with mechanized tools: efficiency goes though the roof. Still one per home.

Human doing high volume manual labor job where there were much more job than single human could handle: number of humans doing the job now is amount of work divided by amount of work human can handle.

Of course we get ambitious, like Panama Canal building ambitious. But even that can’t absorb the previous admin of people doing that kind of work.

Re: Gemini 3.1 Pro

#856

Earlier quoted context omitted.

> "People underrate Google's cost effectiveness so much. Half price of Opus. HALF." Google undercutting/subsidizing it's own prices to bite into Anthropic's market share (whilst selling at a loss) doesn't automatically mean Google is effective.

Everybody is subsidizing their prices. But Flash is 1/8 the cost of sonnet and its not impressive?

> Everybody is subsidizing their prices.

Inference is profitable but model training needs lot of money.

Re: Gemini 3.1 Pro

#857
Ran a bunch of 3D Modeling benchmarks on Gemini 3.1 vs Gemini 3.

Unsurprisingly 3.1 performs a bit better. But surprisingly it costs 2.6x as much ($0.14 vs. $0.37 per 3D Model Generation) and is 2.5x slower (1m 24s vs. 3m 28s).

To me it feels like "lets increase our thinking budget and call it an improved model!"

Re: Gemini 3.1 Pro

#858

Earlier quoted context omitted.

Google might be a mess now, but they have time. OpenAI and Anthropic are on barrowed time, Google has a built in money printer. They just need to outlast the others.

Google is Google. Too much restrictions on the model output. Ask it to create a pentest or let it request a pub key for ssh access and it will refuse.

I was very surprised to find the opposite yesterday. I was asking ChatGPT about firearms and it hit a safeguard ~”I cannot give gun purchasing advice” so I switched to Gemini, and it happily answered the exact copy/paste question

Historically it was the opposite; OpenAI was yolo and Gemini overly cautious to the point of severely limiting utility

Re: Gemini 3.1 Pro

#859

In an attempt to get outside of benchmark gaming I had it make Platypus on a Tricycle. It's not as good as pelican on bicycle. https://www.svgviewer.dev/s/BiRht5hX

that's better than i thought it would be

would love to be able to teleport this thread to, oh, 5 years ago. people would think some sort of alien technology had landed.

Re: Gemini 3.1 Pro

#860
At risk to be unpopular Gemini 3.0 Pro made a huge difference for me when I moved some workflow to Antigravity, especially compared to ChatGPT.

The latest update? I simply don’t care. I am not paid to evaluate models, I am paid to build. Not sure 4 benchmark points are making the difference.

Post reply on HN