Live data from Hacker News

Gemini 3.1 Pro

blog.google

41–50 of 951 posts

Re: Gemini 3.1 Pro

#43
post #11
post #6

blog post is up- https://blog.google/innovation-and-ai/models-and-research/ge... edit: biggest benchmark changes from 3 pro: arc-agi-2 score went from 31.1% -> 77.1% apex-agents score went from 18.4% -> 33.5%

The touted SVG improvements make me excited for animated pelicans.

I just gave it a shot and this is what I got: https://codepen.io/takoid/pen/wBWLOKj

The model thought for over 5 minutes to produce this. It's not quite photorealistic (some parts are definitely "off"), but this is definitely a significant leap in complexity.

Re: Gemini 3.1 Pro

#44
post #37

Surprisingly big jump in ARC-AGI-2 from 31% to 77%, guess there's some RLHF focused on the benchmark given it was previously far behind the competition and is now ahead. Apart from that, the usual predictable gains in coding. Still is a great sweet-spot for performance, speed and cost. Need to hack Claude Code to use their agentic logic+prompts but use Gemini models. I wish Google also updated Flash-lite to 3.0+, wou…

>I wish Google also updated Flash-lite to 3.0+

I hope every day that they have made gains on their diffusion model. As a sub agent it would be insane, as it's compute light and cranks 1000+ tk/s

Re: Gemini 3.1 Pro

#46
Google is terrible at marketing, but this feels like a big step forward.

As per the announcement, Gemini 3.1 Pro score 68.5% on Terminal-Bench 2.0, which makes it the top performer on the Terminus 2 harness [1]. That harness is a "neutral agent scaffold," built by researchers at Terminal-Bench to compare different LLMs in the same standardized setup (same tools, prompts, etc.).

It's also taken top model place on both the Intelligence Index & Coding Index of Artificial Analysis [2], but on their Agentic Index, it's still lagging behind Opus 4.6, GLM-5, Sonnet 4.6, and GPT-5.2.

---

[1] https://www.tbench.ai/leaderboard/terminal-bench/2.0?agents=...

[2] https://artificialanalysis.ai

Re: Gemini 3.1 Pro

#47

Gemini 3 is pretty good, even Flash is very smart for certain things, and fast! BUT it is not good at all at tool calling and agentic workflows, especially compared to the recent two mini-generations of models (Codex 5.2/5.3, the last two versions of Anthropic models), and also fell behind a bit in reasoning. I hope they manage to improve things on that front, because then Flash would be great for many tasks.

You can really notice the tool use problems. They gotta get on that. The agent trend seems real, and powerful. They can't afford to fall behind on it.

Re: Gemini 3.1 Pro

#48
post #22
post #6

blog post is up- https://blog.google/innovation-and-ai/models-and-research/ge... edit: biggest benchmark changes from 3 pro: arc-agi-2 score went from 31.1% -> 77.1% apex-agents score went from 18.4% -> 33.5%

Does the arc-agi-2 score more than doubling in a .1 release indicate benchmark-maxing? Though i dont know what arc-agi-2 actually tests

I assume all the frontier models are benchmaxxing, so it would make sense

Re: Gemini 3.1 Pro

#49
ok , so they are scared that 5.3 (pro) will be released today/tomorrow and blow it out of the water and rushed it while they could still reference 5.2 benchmarks.

Re: Gemini 3.1 Pro

#50

Gemini 3 is pretty good, even Flash is very smart for certain things, and fast! BUT it is not good at all at tool calling and agentic workflows, especially compared to the recent two mini-generations of models (Codex 5.2/5.3, the last two versions of Anthropic models), and also fell behind a bit in reasoning. I hope they manage to improve things on that front, because then Flash would be great for many tasks.

In other words: they just need to motivate their employees while giving in to finance's demands to fire a few thousand every month or so ...

And don't forget, it's not just direct motivation. You can make yourself indispensable by sabotaging or at least not contributing to your colleagues' efforts. Not helping anyone, by the way, is exactly what your managers want you to do. They will decide what happens, thank you very much, and doing anything outside of your org ... well there's a name for that, isn't there? Betrayal, or perhaps death penalty.

Post reply on HN