Live data from Hacker News

Gemini 3.1 Pro

blog.google

111–120 of 951 posts

Re: Gemini 3.1 Pro

#112
post #72
post #52

Pretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.

Great pelican but what’s up with that fish in the basket?

Where else are cycling Pelican's meant to keep their fish?

Re: Gemini 3.1 Pro

#113

Google is terrible at marketing, but this feels like a big step forward. As per the announcement, Gemini 3.1 Pro score 68.5% on Terminal-Bench 2.0, which makes it the top performer on the Terminus 2 harness [1]. That harness is a "neutral agent scaffold," built by researchers at Terminal-Bench to compare different LLMs in the same standardized setup (same tools, prompts, etc.). It's also taken top model place on both…

Benchmarks aren't everything.

Gemini consistently has the best benchmarks but the worst actual real-world results.

Every time they announce the best benchmarks I try again at using their tools and products and each time I immediately go back to Claude and Codex models because Google is just so terrible at building actual products.

They are good at research and benchmaxxing, but the day to day usage of the products and tools is horrible.

Try using Google Antigravity and you will not make it an hour before switching back to Codex or Claude Code, it's so incredibly shitty.

Re: Gemini 3.1 Pro

#115
post #52

Pretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.

Models are soon going to start benchmaxxing generating SVGs of pelicans on bikes

Soon? I'd be willing to bet it's been included in the training set at least 6 months by now. Not so obvious so it generates always perfect pelicans on bikes, but sufficiently for the "minibench" to be less useful today than in the past.

Re: Gemini 3.1 Pro

#116
post #66
post #52

Pretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.

Not even animated? This is 2026.

Jeff Dean just posted an animated version: https://x.com/JeffDean/status/2024525132266688757

Re: Gemini 3.1 Pro

#117

We've gone from yearly releases to quarterly releases. If the pace of releases continues to accelerate - by mid 2027 or 2028 we're headed to weekly releases.

But actual progress seems to be slower. These modes are releasing more often but aren’t big leaps.

Re: Gemini 3.1 Pro

#119
post #52

Pretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.

It's an excellent demonstration of the main issue I have with the Gemini family of models, they always go "above and beyond" to do a lot of stuff, even if I explicitly prompt against it. In this case, most of the SVG ends up consisting not just of a bike and a pelican, but clouds, a sun, a hat on the pelican and so much more. Exactly the same thing happens when you code, it's almost impossible to get Gemini to not do…

Do you have Personalization Instructions set up for your LLM models?

You can make their responses fairly dry/brief.

Re: Gemini 3.1 Pro

#120
post #52

Pretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.

I hope we keep beating this dead horse some more, I'm still not tired of it.
Post reply on HN