Live data from Hacker News

Gemini 3.1 Pro

blog.google

441–450 of 951 posts

Re: Gemini 3.1 Pro

#441
post #52

Pretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.

What's crazy is you've influenced them to spend real effort ensuring their model is good at generating animated svgs of animals operating vehicles. The most absurd benchmaxxing. https://x.com/jeffdean/status/2024525132266688757?s=46&t=ZjF...

You don't have to benchmax everything, just the benchmarks in the right social circles

Re: Gemini 3.1 Pro

#442

Earlier quoted context omitted.

More like half of Google's AI team is hanging out on HN, and they can optimise for that outcome to get a good rep among the dev community.

Hello. (I'm not aware of anyone doing this, but GDM is quite info-siloed these days, so my lack of knowledge is not evidence it's not happening)

Hello.

Please push internally for more reliable tool use across Gemini models. Intelligence is useless if it can't be applied :)

Re: Gemini 3.1 Pro

#444
post #332
post #306

These models are so powerful. It's totally possible to build entire software products in the fraction of the time it took before. But, reading the comments here, the behaviors from one version to another point version (not major version mind you) seem very divergent. It feels like we are now able to manage incredibly smart engineers for a month at the price of a good sushi dinner. But it also feels like you have to b…

I had an interesting experience recently where I ran Opus 4.6 against a problem that o4-mini had previously convinced me wasn't tractable... and Opus 4.6 found me a great solution. https://github.com/simonw/sqlite-chronicle/issues/20 This inspired me to point the latest models at a bunch of my older projects, resulting in a flurry of fixes and unblocks.

I have a codebase (personal project) and every time there is a new Claude Opus model I get it to do a full code review. Never had any breakages in last couple of model updates. Worried one day it just generates a binary and deletes all the code.

Re: Gemini 3.1 Pro

#445

Someone needs to make an actual good benchmark for LLM's that matches real world expectations, theres more to benchmarks than accuracy against a dataset.

this reminds me of that joke of someone saying "it's crazy that we have ten different standards for doing this", and then there're 11 standards

Re: Gemini 3.1 Pro

#446

Earlier quoted context omitted.

Is the thinking token stream obfuscated? Im fully immersed

It's just a summary generated by a really tiny model. I guess it also an ad-hoc way to obfuscate it, yes. In particular they're hiding prompt injections they're dynamically adding sometimes. Actual CoT is hidden and entirely different from that summary. It's not very useful for you as a user, though (neither is the summary).

They hide the CoT because they don't want competitors to train on it

Re: Gemini 3.1 Pro

#447

It got the car wash question perfectly: You are definitely going to have to drive it there—unless you want to put it in neutral and push! While 200 feet is a very short and easy walk, if you walk over there without your car, you won't have anything to wash once you arrive. The car needs to make the trip with you so it can get the soap and water. Since it's basically right next door, it'll be the shortest drive of you…

Some people are suggesting that this might actually be in the training set. Since I can't rule that out, I tried a different version of the question, with an elephant instead of a car: > It's a hot and dusty day in Arizona and I need to wash my elephant. There's a creek 300 feet away. Should I ride my elephant there or should I just walk there by myself? Gemini said: That sounds like quite the dusty predicament! Give…

From Gemini pro:

You should definitely ride the elephant (or at least lead it there)!

Here is the logic:

If you walk there by yourself, you will arrive at the creek, but the dirty elephant will still be 300 feet back where you started. You can't wash the elephant if it isn't with you!

Plus, it is much easier to take the elephant to the water than it is to carry enough buckets of water 300 feet back to the elephant.

Would you like another riddle, or perhaps some actual tips on how to keep cool in the Arizona heat?

Re: Gemini 3.1 Pro

#448

Earlier quoted context omitted.

What's crazy is you've influenced them to spend real effort ensuring their model is good at generating animated svgs of animals operating vehicles. The most absurd benchmaxxing. https://x.com/jeffdean/status/2024525132266688757?s=46&t=ZjF...

So let's put things we're interested in in the benchmarks. I'm not against pelicans!

I think the reason the pelican example is great is because it's bizarre enough that it's unlikely that to appear in the training as one unified picture.

If we picked something more common, like say, a hot dog with toppings, then the training contamination is much harder to control.

Re: Gemini 3.1 Pro

#450
My enthusiasm is a bit muted this cycle because I've been burned by Gemini CLI. These models are very capable but Gemini CLI just doesn't seem to be able to work for one it never follows instructions strictly like its competitors do, and it hallucinates even which is a rarity.

More importantly feels like Google is stretched thin across different Gemini products and pricing reflects this, I still have no idea how to pay for Gemini CLI, in codex/claude its very simple $20/month for entry and $200/month for ton of weekly usage.

I hope whoever is reading this from Google they can redeem Gemini CLI by focusing on being competitive instead of making it look pretty (that seems to be the impression I got from the updates on X)

Post reply on HN