Live data from Hacker News

Gemini 3.1 Pro

blog.google

241–250 of 951 posts

Re: Gemini 3.1 Pro

#241
post #52

Pretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.

It seems they trained the model to output good svg’s.

In their blog post[1], first use case they mention is svg generation. Thus, it might not be any indicator at all anymore.

[1] https://blog.google/innovation-and-ai/models-and-research/ge...

Re: Gemini 3.1 Pro

#242
post #52

Pretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.

Does anyone understand why LLMs have gotten so good at this? Their ability to generate accurate SVG shapes seems to greatly outshine what I would expect, given their mediocre spatial understanding in other contexts.

Re: Gemini 3.1 Pro

#243

Earlier quoted context omitted.

Simons been doing this exact test for nearly 18 months now, if vendors want to benchmaxx it then they've had more than enough time to do so already.

Exactly. As far as I'm concerned, the benchmark is useless. It's way too easy and rewarding to train on it.

I mean if you want to make your own benchmark, simply don't make it public and don't do it often. If your salamander on skis or whatever gets better with time it likely has nothing to do with being benchmaxxed.

Re: Gemini 3.1 Pro

#244

Implementation and Sustainability Hardware: Gemini 3 Pro was trained using Google’s Tensor Processing Units (TPUs). TPUs are specically designed to handle the massive computations involved in training LLMs and can speed up training considerably compared to CPUs. TPUs often come with large amounts of high-bandwidth memory, allowing for the handling of large models and batch sizes during training, which can lead to bet…

Bla bla bla yada sustainability yada often come with large better growing faster...

It's such an uninformative piece of marketing crap

Re: Gemini 3.1 Pro

#245

Earlier quoted context omitted.

Ugh, the gears and chain don't mesh and there's no sprocket on the rear hub But seriously, I can't believe LLMs are able to one-shot a pelican on a bicycle this well. I wouldn't have guessed this was going to emerge as a capability from LLMs 6 years ago. I see why it does now, but... It still amazes me that they're so good at some things.

next time you host a party, have people try to draw a bicycle on your whiteboard (you have a whiteboard in your house right? you should, anyway...) human adults are generally quite bad at drawing them, unless they spend a lot of time actually thinking about bicycles as objects

They are, and it is very funny.

https://www.behance.net/gallery/35437979/Velocipedia

Re: Gemini 3.1 Pro

#246
post #52

Pretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.

At this point, the pelican benchmark became so widely used that there must be high quality pelicans in the dataset, I presume. What about generating an okapi on a bicycle instead?

Re: Gemini 3.1 Pro

#247

Price is unchanged from Gemini 3 Pro: $2/M input, $12/M output. https://ai.google.dev/gemini-api/docs/pricing Knowledge cutoff is unchanged at Jan 2025. Gemini 3.1 Pro supports "medium" thinking where Gemini 3 did not: https://ai.google.dev/gemini-api/docs/gemini-3 Compare to Opus 4.6's $5/M input, $25/M output. If Gemini 3.1 Pro does indeed have similar performance, the price difference is notable.

If we don't see a huge gain on the long-term horizon thinking reflected with the Vendor-Bench 2, I'm not going to switch away from CC. Until Google can beat Anthropic on that front, Claude Code paired with the top long-horizon models will continue to pull away with full stack optimizations at every layer.

Re: Gemini 3.1 Pro

#248

Earlier quoted context omitted.

Less pretty and more practical, it's really good at outputting circuit designs as SVG schematics. https://www.svgviewer.dev/s/dEdbH8Sw

I don't know what of this is the prompt and what was the output, but that's a pretty bad schematic (for both aesthetic and circuit-design reasons).

Yes but you concede it is a schematic.

Re: Gemini 3.1 Pro

#249

Earlier quoted context omitted.

Jeff Dean just posted an animated version: https://x.com/JeffDean/status/2024525132266688757

One underrated thing about the recent frontier models, IMO, is that they are obviating the need for image gen as a standalone thing. Opus 4.6 (and apparently 3.1 Pro as well) doesn't have the ability to generate images but it is so good at making SVG that it basically doesn't matter at this point. And the benefit of SVG is that it can be animated and interactive. I find this fascinating because it literally just happ…

2025 that is
Post reply on HN