Live data from Hacker News

Gemini 3.1 Pro

blog.google

351–360 of 951 posts

Re: Gemini 3.1 Pro

#351
post #72
post #52

Pretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.

Great pelican but what’s up with that fish in the basket?

Yeah, why only _one_ fish?

It's obvious that pelican is riding long distance, no way a single fish is sufficiently energy dense for more than a few miles.

Can't the model do basic math???

Re: Gemini 3.1 Pro

#352
post #11
post #6

blog post is up- https://blog.google/innovation-and-ai/models-and-research/ge... edit: biggest benchmark changes from 3 pro: arc-agi-2 score went from 31.1% -> 77.1% apex-agents score went from 18.4% -> 33.5%

The touted SVG improvements make me excited for animated pelicans.

My perennial joke is as soon as that got on HN front page Google went and hired some interns and they spend a 100% of the time on pelicans.

Re: Gemini 3.1 Pro

#353

Earlier quoted context omitted.

What. You don't have yours ask for edit approval?

Who has time for that? This is how I run codex: `codex --sandbox danger-full-access --dangerously-bypass-approvals-and-sandbox --search exec "$PROMPT"`, having to approve each change would effectively destroy the entire point of using an agent, at least for me. Edit: obviously inside something so it doesn't have access to the rest of my system, but enough access to be useful.

I wouldn't even think of letting an agent work in that made. Even the best of them produce garbage code unless I keep them on a tight leash. And no, not a skill issue.

What I don't have time to do is debug obvious slop.

Re: Gemini 3.1 Pro

#355

Every time I've used Gemini models for anything besides code or agentic work they lean so far into the RLHF induced bold lettering and bullet point list barf that everything they output reads as if the model was talking _at_ me and not _with_ me. In my Openclaw experiment(s) and in the Gemini web UI, I've specifically added instructions to avoid this type of behavior, but it only seemed to obey those rules when I rem…

It definitely has the worst "voice" in my opinion. Feels very overachieving McKinsey intern to me.

Re: Gemini 3.1 Pro

#356
The speed of these 3.1 and Preview releases is starting to feel like the early days of web frameworks. It’s becoming less about the raw benchmarks and more about which model handles long-context 'hallucination' well enough to be actually used in a production pipeline without constant babysitting.

Re: Gemini 3.1 Pro

#357

Seems like they actually fixed some of the problems with the model. Hallucinations rate seems to be much better. Seems like they also tuned the reasoning maybe that were they got most of the improvements from.

The hallucination rate with the Gemini family has always been my problem with them. Over the last year they’ve made a lot of progress catching the Gemini models up to/near the frontier in general capability and intelligence, but they still felt very late 2024 in terms of hallucination rate. Which made the Gemini models untrustworthy for anything remotely serious, at least in my eyes. If they’ve fixed this or at least…

Maybe I haven't kept up with how ghatgpt and claude are doing , but 6 monthlatelys ago or so, I thought Gemini was leading on that front.

Re: Gemini 3.1 Pro

#358
post #334
post #14

Earlier quoted context omitted.

> increasing the number for such a minor change is not a move in the right direction A .1 model number increase seems reasonable for more than doubling ARC-AGI 2 score and increasing so many other benchmarks. What would you have named it?

My issue is that we haven't even gotten the release version of 3.0, that is also still in Preview, so may stick with 3.0 till that has been deemed stable. Basically, what does the word "Preview" mean, if newer releases happen before a Preview model is stable? In prior Google models, Preview meant that there'd still be updates and improvements to said model prior to full deployment, something we saw with 2.5. Now, the…

Given the pace AI is improving and that it doesn't give the exact same answers under many circumstances, is the the [in]stability of "preview" a concern?

GMail was in "beta" for 5 years.

Re: Gemini 3.1 Pro

#359
post #65

Earlier quoted context omitted.

I'm thinking now that as models get better and better at generating SVGs, there could be a point where we can use them to just make arbitrary UIs and interactive media with raw SVGs in realtime (like flash games).

Or quite literally a game where SVG assets are generated on the fly using this model

Thats one dimension before another long term milestone: Realtime generation of 3D mesh content during gameplay.

Which is the "left brain" approach vs the "right brain" approach of coming at dynamic videogames from the diffusion model direction which the Gemini Genie thing seems to be about.

Re: Gemini 3.1 Pro

#360

I really want to use google’s models but they have the classic Google product problem that we all like to complain about. I am legit scared to login and use Gemini CLI because the last time I thought I was using my “free” account allowance via Google workspace. Ended up spending $10 before realizing it was API billing and the UI was so hard to figure out I gave up. I’m sure I can spend 20-40 more mins to sort this ou…

I've been using it lately with OpenCode and it's working pretty well (except for API reliability issues).
Post reply on HN