Live data from Hacker News

Gemini 3.1 Pro

blog.google

501–510 of 951 posts

Re: Gemini 3.1 Pro

#501

Earlier quoted context omitted.

Models are soon going to start benchmaxxing generating SVGs of pelicans on bikes

Soon? I'd be willing to bet it's been included in the training set at least 6 months by now. Not so obvious so it generates always perfect pelicans on bikes, but sufficiently for the "minibench" to be less useful today than in the past.

If only there were some way to test it, like swapping the two nouns in the sentence. Alas.

Re: Gemini 3.1 Pro

#502

Earlier quoted context omitted.

Is the thinking token stream obfuscated? Im fully immersed

It's just a summary generated by a really tiny model. I guess it also an ad-hoc way to obfuscate it, yes. In particular they're hiding prompt injections they're dynamically adding sometimes. Actual CoT is hidden and entirely different from that summary. It's not very useful for you as a user, though (neither is the summary).

The early version of Gemini 2.5 did initially show the actual CoT in AI Studio, and it was pretty interesting in some cases.

Re: Gemini 3.1 Pro

#504
post #376

Gets 10/10 on my potato benchmarks: https://aibenchy.com/model/google-gemini-3-1-pro-preview-med...

Added one more test, which surprisingly gemini flash 3 reasoning passes, but gemini 3.1 pro not

Re: Gemini 3.1 Pro

#505

Can we switch from Claude Code to Google yet? Benchmarks are saying: just try But real world could be different

My sense is that the Gemini models are very capable but the Gemini CLI experience is subpar compared to Claude Code and Codex. I'm guess that it's the harness but since it can get confused, fall into doom loops, and generally lose the plot in a way that the model does not in Gemini Studio or the Gemini app. I think a bunch of these harnesses are open source so it surprises me that there can be such a gulf between the…

Its not just subpar, its not even sub-sub-par.

It goes into loops and never completes a task 8 times out of 10 that i've used it.

Re: Gemini 3.1 Pro

#507

Earlier quoted context omitted.

Good to see it wearing a helmet. Their safety team must be on their game.

Yes but why would a pelican need a helmet? If it falls over it can just fly away... Common sense 1 Gemini 0

Why would a pelican be riding a bicycle at all, for that matter?

Re: Gemini 3.1 Pro

#508
post #258

Earlier quoted context omitted.

This feels very Google

I found the Googler!

Nope. The closest I've gotten was rejecting Google recruiters several times.

But like everyone else I'm used to Google failing to care about products.

Re: Gemini 3.1 Pro

#509

Earlier quoted context omitted.

What's crazy is you've influenced them to spend real effort ensuring their model is good at generating animated svgs of animals operating vehicles. The most absurd benchmaxxing. https://x.com/jeffdean/status/2024525132266688757?s=46&t=ZjF...

Animated SVG is huge. People in different professions are worrying to different degrees in terms of being replaced by ML, but this one is huge with regards to digital art.

yeah, complex SVG's are so much more bandwidth, computation and energy efficient than raster images - up to a point! but in general use we are not at that point and there's so much more we can do with it

I've been meaning to let coding agents take a stab at using the lottie library https://github.com/airbnb/lottie-web to supercharge the user experience without needing to make it a full time job

Re: Gemini 3.1 Pro

#510

Earlier quoted context omitted.

your question may have become part of the training data with how much coverage there was around it. perhaps you should devise a new test :P

3.1 Pro has the same Jan 2025 knowledge cutoff as the other 3 series models. So if 3.1 has it in its training data, the other ones would have as well.

The fact it's still Jan 2025 is weird to me. Have they not have a successful pretrain in over a year?
Post reply on HN