Live data from Hacker News

Gemini 3.1 Pro

blog.google

901–910 of 951 posts

Re: Gemini 3.1 Pro

#901
post #751

Earlier quoted context omitted.

Strange that you say that because the general consensus (and my experience) seems to be the opposite, as well as the AA-Omniscience Hallucination Rate Benchmark which puts 3.0 Pro among the higher hallucinating models. 3.1 seems to be a noticeable improvement though.

Google actually has the BEST ratings in the AA-Omniscience Index: AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Gemini 3.1 is the top spot, followed by 3.0 and then opus 4.6 max

This isn't actually correct.

Gemini 3.0 gets a very high score because it's very often correct, but it does not have a low hallucination rate.

https://artificialanalysis.ai/#aa-omniscience-hallucination-...

It looks like 3.1 is a big improvement in this regard, it hallucinates a lot less.

Re: Gemini 3.1 Pro

#902
post #751

Earlier quoted context omitted.

Strange that you say that because the general consensus (and my experience) seems to be the opposite, as well as the AA-Omniscience Hallucination Rate Benchmark which puts 3.0 Pro among the higher hallucinating models. 3.1 seems to be a noticeable improvement though.

> the AA-Omniscience Hallucination Rate Benchmark which puts 3.0 Pro among the higher hallucinating models. 3.1 seems to be a noticeable improvement though. As sibling comment says, AA-Omniscience Hallucination Rate Benchmark puts Gemini 3.0 as the best performing aside from Gemini 3.1 preview. https://artificialanalysis.ai/evaluations/omniscience

You are misreading the benchmark.

https://artificialanalysis.ai/#aa-omniscience-hallucination-...

If you look at the results 3.0 hallucinates an awful lot, when it's wrong.

It's just not wrong that often.

(And it looks like 3.1 does better on both fronts)

Re: Gemini 3.1 Pro

#903
post #874

Earlier quoted context omitted.

ChatGPT 4.5 was never released to the public, but it is widely believed to be the foundation the 5.x series is built on. Wonder how GP feels about the minor bumps for other model providers?

Minor version bumps are good and I want model providers to communicate changes. The issue I am having is that Gemini "preview" class models have different deprecation timelines and rate limits, making them impossible to rely on for professional use cases. That's why I'd prefer they finish the 3.0 role out prior to putting resources into deploying a second "preview" class model. For a stable deployment, Google needs a…

Sorry, but you come off as an armchair devops saying things like this. Google is fine, they know more than anyone else about how to run Ai at scale.

"preview" != GA, sounds like you need to adjust your expectations

Re: Gemini 3.1 Pro

#904
post #537

Earlier quoted context omitted.

Thinking is just tacked on for Anthropic's models and always has been so leaving it off actually produces better results everytime.

What about for analysis/planning? Honestly I've been using thinking, but if I don't have to with Opus 4.6 I'm totally keen to turn it off. Faster is better.

I've always just used the "Plan mode" in Claude Code, I don't know if it uses thinking? I have "MAX_THINKING_TOKENS" in my settings.json set to "0", too. Didn't notice a drop in performance, I find it better because it doesn't overthink ("wait, let me try..."). Likely depends on a case-by-case basis (as so often with AI). For me, it's better without thinking.

Re: Gemini 3.1 Pro

#905

If it’s any consolation, it was able to one-shot a UI & data sync race condition that even Opus 4.6 struggled to fix (across 3 attempts). So far I like how it’s less verbose than its predecessor. Seems to get to the point quicker too. While it gives me hope, I am going to play it by the ear. Otherwise it’s going to be - Gemini for world knowledge/general intelligence/R&D and Opus/Sonnet 4.6 to finish it off. UPDATE:…

Interesting, I've had similar issues. It seems to be very clumsy when using its internal tooling. I've seen diffs where it accidentally garbled significant amounts of code, which it then had to go in and manually fix. It's also introduced bugs into features that it wasn't supposed to be touching, and when I asked it why it was making changes to I the other code, it answered that it had failed to copy-paste since larg…

Yeah, I whole heartedly agree with this. Even Codex does this sometimes, although it has been consistently much better than the others at following instructions.

The problem is again that you can’t ever fully trust an agent did exactly what you asked for and in the exact manner that you had hoped.

It works just like you’re dealing with a human companion. Trust takes time to build. Over the period you realize the other individuals weaknesses and support them there.

What makes it a bit challenging right now is the pace of innovation. By the time we get used to a model’s personality, a new update comes out that alters it in unknown ways. Now you’re back to square one.

I’ve been experimenting with asking one frontier model to check on another’s work. That’s proven to be better than doing nothing. Usually they’ll have some genuinely useful feedback.

Re: Gemini 3.1 Pro

#906

Earlier quoted context omitted.

Gemini is the most paradoxical model because it benchmarks great even in private benchmarks done by regular people, Deep Mind is unquestionably full of capable engineers with incredible skill, and personally Gemini has been great for my day job and my coding for fun (not for profit) endeavors. Switching between it and 4.6 in antigravity and I don't see much of a difference, they both do what I ask. But man, people ar…

People can be and often are wrong. You'd notice how good Opus is in Claude Code. IMHO CC is the secret sauce

Opus is just as good in pi.dev, Amp, or OpenCode. CC is an increasingly bug ridden slopfest.

Re: Gemini 3.1 Pro

#907
post #895

Earlier quoted context omitted.

https://arcprize.org/arc-agi/1/ It's a sort of arbitrary pattern matching thing that can't be trained on in the sense that the MMLU can be, but you can definitely generate billions of examples of this kind of task and train on it, and it will not make the model better on any other task. So in that sense, it absolutely can be. I think it's been harder to solve because it's a visual puzzle, and we know how well today's…

The real question is: Why are people designing benchmarks that, if a model is trained on them, it won't improve the performance of the model at any real-world tasks? Why would anyone care about such benchmarks?

People are like typewriter monkeys, if something is possible to make it'll eventually be made.

Re: Gemini 3.1 Pro

#908

Earlier quoted context omitted.

Yes, this is very true and it speaks strongly to this wayward notion of 'models' - it depends so much on the tuning, the harness, the tools. I think it speaks to the broader notion of AGI as well. Claude is definitively trained on the process of coding not just the code, that much is clear. Codex has the same limitation but not quite as bad. This may be a result of Anthropic using 'user cues' with respect to what are…

I know this is only a partial answer, but I feel like Google is once again trying to build a product based on internal priorities, existing business protectionism, and internal business goals, rather than building a product that is listening actively to real use feedback as the primary priority. It is the company’s constant kryptonite. They seem to be, from my third part perspective, repeating the same ol’, same ol’…

What do you think Microsoft is doing? :)

Re: Gemini 3.1 Pro

#909
post #822

I'm doing Ruby and Gemini 3.0 pro has by far been the best model for me. It writes the nicest ruby code, like I would. Further, it either succeeds or fails hard and obviously. I prefer it failing hard instead of of slowly going weird in my code. Similar in antigravity. Privately it's my absolute favorite. So I'm actually rooting for this.

Which harness? Gemini CLI or OpenCode?

Re: Gemini 3.1 Pro

#910
post #709

What I’m noticing, overall: I’ve never cut so much code in my life. I’ve become a coding monster with one of those dark green GitHub profiles ever since 5.3-Codex gave me the confidence to load in a ridiculous number of tasks every day and let it rip. I have about three coding tasks going at once and in another window, Claude Cowork is ripping through PowerPoints and getting back to lawyers. This tech is not going to…

How do you give it tasks? As GitHub issues?
Post reply on HN